\

DeepSeek V4 Pro 0813

267 points - today at 4:04 PM

Source
  • scrlk

    today at 4:42 PM

    Benchmarks:

        | Benchmark                | DS-V4-Pro | DS-V4-Flash | DS-V4-Pro | DS-V4-Flash | GLM-5.2   | Kimi-K3   | Opus-4.8  | Fable 5       |
        |                          | 0813      | 0731        | Preview   | Preview     |           |           |           | (w/ fallback) |
        |--------------------------|-----------|-------------|-----------|-------------|-----------|-----------|-----------|---------------|
        | HLE (wo/w tools)         | 42.7/60.0 | 37.8/51.5   | 37.7/48.2 | 34.8/45.1   | 40.5/54.7 | 43.5/56.0 | 49.8/57.9 | 53.3/63.0     |
        | Terminal Bench 2.1       | 87.9      | 82.7        | 72.1      | 61.8        | 81.0      | 88.3      | 85.0      | 88.0          |
        | NL2Repo                  | 61.5      | 54.2        | 38.5      | 39.4        | 48.9      | -         | 69.7      | -             |
        | Cybergym                 | 83.3      | 76.7        | 52.7      | 38.7        | -         | 80.0      | 78.3      | 83.1          |
        | DeepSWE                  | 62.7      | 54.4        | 12.8      | 7.3         | 46.2      | 67.5      | 58.0      | 70.0          |
        | Toolathlon-Verified      | 74.1      | 70.3        | 55.9      | 49.7        | 59.9      | 76.5      | 76.2      | 77.9          |
        | Agents' Last Exam        | 25.7      | 25.2        | 16.5      | 15.8        | 23.8      | 27.6      | 25.7      | -             |
        | AutomationBench (Public) | 31.8      | 25.1        | 12.8      | 10.8        | 12.9      | 30.8      | 27.2      | 29.1          |
        | DSBench-FullStack        | 71.1      | 68.7        | 41.8      | 37.0        | 61.8      | 73.7      | 71.6      | 77.2          |
        | DSBench-Hard             | 67.2      | 59.6        | 31.1      | 25.8        | 54.5      | 63.0      | 71.7      | 68.3          |
    
    Source: https://reddit.com/r/LocalLLaMA/comments/1vmi0fg/deepseek_v4...

      • parsimo2010

        today at 5:16 PM

        The timing looks like they are trying to take the wind out of Qwen's sails by releasing this on the same day that Qwen released the weights of Qwen3.8-max. Or maybe it's coincidence...

        For comparison I looked at Qwen's claimed benchmarks for Qwen3.8-max (https://qwen.ai/blog?id=qwen3.8). Assuming each published set of benchmarks is believable, it looks like v4 Pro 0813 is better on average but overall performance is comparable. Pro 0813 is much cheaper. If you don't need vision capabilities then you don't have much reason to use Qwen3.8-max.

        - 43.6 on HLE (Presumably without tools). Pro 0813 is a little worse.

        - 86.6 on Terminal Bench 2.1. Pro 0813 is better.

        - 55.9 on NL2Repo. Pro 0813 is better.

        - 27 on Agent's Last Exam. Pro 0813 is a little worse.

        - 72.5 on Toolathon-Verified. Pro 0813 is better.

        - 56.6 on DeepSWE 1.1. If the DeepSWE listed for Pro 0813 is the same version, then Pro is better.

        - 27.3 on AutomationBench. If the AutomationBench (Public) listed for Pro 0813 is the same, then Pro is better.

        I guess we do need to wait to see if the upcoming DS pricing increase is enough to change the value proposition. As it is now, they could double or triple prices and it still would be a better value to use DS. I bet they know that.

          • trollbridge

            today at 5:23 PM

            By that standard, the release of Grok 4.6 was also timed on the same day.

            Given how I think DeepSeek operates... I think they just release it when they feel it's ready, and don't even seem that concerned with what other people are doing.

              • somenameforme

                today at 5:28 PM

                Their leaks would confirm this sort of attitude. They're not trying to become the top player or anything like that - just working to play their part in pushing LLM tech forward and going from there. It was quite refreshing from the 'here's how we're going to dominate the world' nonsense. It's undoubtedly the same attitude that just lets them shrug and cancel the fund raising round after the leaks came from said funding round.

                  • trollbridge

                    today at 5:36 PM

                    The founder of DS's stated goal is to get to AGI. He thinks this is the path to get there.

                    Kind of interesting, when compared to the hubris from American frontier labs.

                    • scrlk

                      today at 5:35 PM

                      Benefits of having a well performing hedge fund funding DeepSeek.

                      IIRC, Demis attempted to start a fund inside DeepMind but it was killed off. In an alternative world where he manages to pull that off, perhaps DeepMind would still be independent with Demis at the helm.

                      • surgical_fire

                        today at 5:33 PM

                        Their stance on LLM development is why they earned my respect in a time when OpenAI and Anthropic only earn my mistrust.

                        That, and the fact that DS is an insanely capable model.

                    • parsimo2010

                      today at 5:32 PM

                      Actually, yes. I just didn't know about Grok's release because they aren't on the front page of HN.

                  • eli

                    today at 5:24 PM

                    Official pricing only kinda matters for an open weight model, no?

                      • parsimo2010

                        today at 5:31 PM

                        It still matters as a point of comparison until other providers come online. If the consensus price from other providers is much different that can be compared then. But for now we have $0.435 / $0.87 for v4 Pro 0813 (with increase announced but we don't know the new pricing), and $2 / $6 for Qwen3.8-max. So until we get other data points that is what we have to look at.

                          • eli

                            today at 5:39 PM

                            I wondered if the promised change in pricing is actually going to be deepseek bringing up their cached costs. They're extremely inexpensive.

                    • maherbeg

                      today at 5:26 PM

                      I mean at the rate of model releases happening, I think a lot of these will collide more often than expected!

                  • goldenarm

                    today at 5:19 PM

                    Geometric mean of all these benchmarks :

                    * GPT-5.6 Sol: 65.5

                    * Fable 5 (w/ fallback): 64.5

                    * Opus 5: 64.0

                    * DS-V4-Pro 0813: 62.5

                    * Kimi-K3: 62.3

                    * DS-V4-Flash 0731: 55.8

                    * GLM-5.2: 47.3

                    • bel8

                      today at 5:17 PM

                      So it's a Fable class LLM?

                                                   DSV4Pro vs Fable5
                          HLE w tools              60.0 vs 63.0
                          Terminal Bench 2.1       87.9 vs 88.0
                          Cybergym                 83.3 vs 83.1
                          DeepSWE                  62.7 vs 70.0
                          Toolathlon-Verified      74.1 vs 77.9
                          AutomationBench (Public) 31.8 vs 29.1
                          DSBench-FullStack        71.1 vs 77.2
                          DSBench-Hard             67.2 vs 68.3

                        • eli

                          today at 5:20 PM

                          Fable's guardrails would never let it do something like Cybergym so at least for that one it's measuring Opus 5

                            • wren6991

                              today at 5:26 PM

                              We have a first-party figure from the system card [1]:

                              > Mythos 5 reproduced 83.8% of targeted vulnerabilities on a single try, and produced at least one crash in 99.4% of tasks. This is comparable to Claude Mythos Preview, which reproduced 83.1% of targeted vulnerabilities and produced a crash in 97.1% of tasks. By contrast, Claude Opus 4.8 achieved a score of 78.1% (95.7% any crash).

                              So their quoted figure exactly matches the figure for Mythos Preview, although they don't state the provenance. It could also quite possibly be an independent measurement of Opus 5.

                              [1]: https://www-cdn.anthropic.com/57a52ea7d8f0e54e8a542e90826608...

                          • aftbit

                            today at 5:19 PM

                            Fabble lol

                              • qiran87

                                today at 5:53 PM

                                [dead]

                    • jklmnopqrstuvw

                      today at 5:34 PM

                      Tested both DS v4 pro 0813 and Grok 4.6 (all from openrouter) on Codex cli. Worked on a same new feature development on my project.

                      Deepseek 4 pro: Worked for 12m 02s - cost $0.12 - has bug.

                      Grok 4.6: Worked for 3m 18s - cost $ 1.41 - no bug.

                      BTW, why grok 4.6 news being down voted and disappeared from frontpage?

                        • computerex

                          today at 5:39 PM

                          Repeat the test like 5 times for each model and see the results.

                          • ferongr

                            today at 5:35 PM

                            Rocket man bad.

                              • nozzlegear

                                today at 5:41 PM

                                This but unironically

                                  • today at 5:52 PM

                                • today at 5:42 PM

                          • alecsm

                            today at 5:24 PM

                            I've been using the last Deepseek Flash update for a week and I'm amazed. It was a capable model for easy tasks but now it looks like it can do some heavy development for peanuts.

                            I can't wait to try this new one.

                            • aabdi

                              today at 4:06 PM

                              https://api-docs.deepseek.com/quick_start/pricing/

                              Competitive with opus 4.8 but weaker than sol or fable. About 20x cheaper.

                                • xynelius

                                  today at 5:20 PM

                                  If that wasn't impressive enough, it's actually ~60x cheaper if you take into account the typical cache-read/input/output split in agentic coding, and the deep discount for cache reads offered by DeepSeek. Opencode has some public data on the typical split [1]:

                                  For DeepSeek V4 Pro the typical split is 750 in, 290 out, 82k cached.

                                  Cost per request for V4 Pro: $0.000875 per request.

                                  Equivalent Opus cost (w/o taking into account cache write costs): $0.052 per request.

                                  [1] https://opencode.ai/docs/go/#usage-limits

                                  • JacobAsmuth

                                    today at 5:06 PM

                                    Per token. You need to look at pricing per task.

                                      • trollbridge

                                        today at 5:25 PM

                                        ... which still comes out cheaper, since DeepSeek caches so much more.

                                        I keep track of my token consumption even on subscription plans and my equiv. cost for my 5.6-Sol usage is around $4000-$8000 a month.

                                    • swiftcoder

                                      today at 4:36 PM

                                      How does it stack against the updated Deepseek Flash version?

                                        • pixelesque

                                          today at 4:57 PM

                                          I've found Pro to be a lot better per "task" than the recently released Flash for code reviews and things (via OpenRouter running in pi.dev).

                                          Flash makes a lot more initial mistakes, and then has to re-check stuff, and produces much more output compared to Pro. It often gets to the correct result eventually, but the output volume is often 5x more than for Pro, and the initial outputs are often wrong, with the first few saying something wrong (like there's a bug, or the code won't compile when it does), and then saying things like "Wait, let me re-check:", or "Actually, looking at it more carefully:" and then it thinks a bit more and eventually gets to the right answer.

                                            • surgical_fire

                                              today at 5:17 PM

                                              I use a plan -> implement wotkflow for this reason.

                                              pro plans, flash implements. I am super happy with how flash behaves like that.

                                              • swiftcoder

                                                today at 5:00 PM

                                                yeah, I've definitely noticed one has to be quite precise to keep Flash on the straight-and-narrow

                                            • k__

                                              today at 4:42 PM

                                              Around 5 percentage points better. (E.g., 87% instead of 82%)

                                                • Gecko4072

                                                  today at 4:46 PM

                                                  So not worth it over flash? Even at ~7x the size it isn't worth the price hike. Flash may be a monster of a model due to all the RL it received from free usage everywhere.

                                                    • networked

                                                      today at 5:00 PM

                                                      I haven't tried DeepSeek V4 Pro 0813 yet. Recent experience tells me that larger models are worth it in non-obvious ways. MiMo-V2.5-Pro solved problems that DeepSeek V4 Flash 0731 couldn't solve for me: for example, adding a live counter for elided reasoning lines to a terminal-based coding harness. You wouldn't be able to tell from the scores on their respective Artifical Analysis page (https://artificialanalysis.ai/models/mimo-v2-5-pro, https://artificialanalysis.ai/models/deepseek-v4-flash). I like the DeepSeek V4 models, though. They critiqued my engineering decisions better than MiMo, and they seem to have a distinct aesthetic in the SVGs they write.

                                                        • trollbridge

                                                          today at 5:27 PM

                                                          Interesting - I've been dropping into MiMo-V2.5-Pro-UltraSpeed whenever Flash seems to be "stuck" and it usually figures it out. I use UltraSpeed just because I'm so frustrated by then that I'm impatient.

                                                          I still find 5.6-Sol can solve some things neither of those can, but it's so slow (and it's so hard to trace / debug the reasoning) that I just let it run overnight.

                                                            • networked

                                                              today at 5:43 PM

                                                              What about 5.6 Terra and especially Luna? Luna scores pretty high on benchmarks and seems to have different habits (like a denser pattern of tool use) and blind spots.

                                                              I'm trying out a development workflow where I generate mundane code with MiMo and Luna (and soon V4 Pro 0813?) and have Opus 5 (running on only a Pro subscription) review and refactor it. I'm not sure it will justify the context switching, but it's an interesting exercise.

                                                      • saaga

                                                        today at 4:53 PM

                                                        Yea that's what I was thinking. Flash is nuts. I find I have to be a more precise and specific with it but damn. It's crossed a threshold of production grade coding for sure.

                                                        I was running a session over a couple days and it didnt cross a dollar lol.

                                                        • eli

                                                          today at 5:29 PM

                                                          Opus 5 medium to Opus 5 max is only 3 points, if that puts it in context

                                                          • npn

                                                            today at 5:10 PM

                                                            I still believe this is not the full potential of pro models. I expect they will release another checkpoint later this year.

                                                            • k__

                                                              today at 4:49 PM

                                                              I tried the previous Pro model and in the end it was 50% more expensive than the previous Flash.

                                                              Wasn't worth it.

                                                          • sparkling

                                                            today at 4:51 PM

                                                            deepseek-v4-flash feels so fast and snappy, i'm loving it. Happy to trade speed for the the 5% degraded benchmarking performance.

                                                              • saaga

                                                                today at 4:53 PM

                                                                I feel the same too. I like the speed. I'm also a big fan of glm 5.2 fast. I can't wait for like 2000 t/s on these haha.

                                                                • k__

                                                                  today at 4:53 PM

                                                                  I wouldn't exactly call it snappy, but faster than Pro, yes.

                                                                    • ericd

                                                                      today at 4:57 PM

                                                                      Single request depth on vllm with dspark, I'm getting ~200 tps, I'd say it's pretty snappy.

                                                                        • JacobAsmuth

                                                                          today at 5:07 PM

                                                                          Well sure but you're running on tens of thousands of dollars of hardware.

                                                  • eshack94

                                                    today at 5:40 PM

                                                    It appears that the only available endpoint (as of this writing) requires enabling "Allow paid endpoints that train on request data" in the OpenRouter privacy settings. I hope additional paid providers will become available that don't require training on data.

                                                      • jubilanti

                                                        today at 5:53 PM

                                                        Their privacy policy doesn't forbid them from just straight up publishing your raw prompts as training data.

                                                        My threat model is that anything I POST to DeepSeek I treat as public to the web, as much as a public GitHub repo is.

                                                        • cdolan

                                                          today at 5:48 PM

                                                          That is likely because Deepseek themselves is the only host.

                                                          In 24-48 hours there will be other options I presume

                                                      • Gecko4072

                                                        today at 4:48 PM

                                                        Currently burning money quickly on official deepseek api. They are also increasing pricing starting today. V4 Flash 0731 still feels like the most outstanding model of the past few months and probably to come.

                                                          • nolist_policy

                                                            today at 5:09 PM

                                                            DeepSeek V4 Flash is the "too cheap to meter" of AI. And you can run the full unquantized model locally for $8000 (2x DGX Spark) at full 1M context and decent speeds: https://github.com/elsung/dgx-spark-deepseek-v4-flash#-long-...

                                                            • sschueller

                                                              today at 5:48 PM

                                                              Deepseek seems to have gotten too cheap. I have been using it for a long time and it's at a point now where my credits balance barely moves even at max setting.

                                                              • eli

                                                                today at 5:33 PM

                                                                The Deepseek official API is good with excellent caching.

                                                                But their privacy policy is unusually bad - they can train off your prompts and completions.

                                                                • Eueudhsbsj32

                                                                  today at 5:12 PM

                                                                  What's the new pricing?

                                                                  The prices on OpenRouter still look the same.

                                                                  • igravious

                                                                    today at 5:15 PM

                                                                    yup :)

                                                                    i'm doing opencode <-> openrouter <-> official deepseek api (i don't get the opencode hate, i like it)

                                                                    how are you doing it?

                                                                    am also using Kimi K3 via kimi-code

                                                                    and also GLM 5.2 via ZCode

                                                                    happy with all three, they're trailing frontier but i figure if i'm running GNU/Linux then i ought to favour open weights models with my €s -- reduced my usage of claude/gpt to the ~$20 tier just to keep abreast of claude_code/codex developments

                                                                    • Jsttan

                                                                      today at 5:07 PM

                                                                      What is the new price through?

                                                                        • Gecko4072

                                                                          today at 5:10 PM

                                                                          https://api-docs.deepseek.com/quick_start/pricing/

                                                                          edit: there are banner announcements saying v4 flash pricing will increase first then overall by an undetermined amount

                                                                            • today at 5:22 PM

                                                                              • nchmy

                                                                                today at 5:13 PM

                                                                                i dont see any price increase there... what am i missing?

                                                                                  • vdfs

                                                                                    today at 5:31 PM

                                                                                    It's a big confusion, some[0] say an email was sent about significant price increase, personal I haven't seen anything official

                                                                                    [0] https://finance.yahoo.com/technology/ai/articles/deepseek-pl...

                                                                                      • surgical_fire

                                                                                        today at 5:35 PM

                                                                                        The email is real, I received it from DeepSeek itself. I probably received it because I buy tokens directly from them.

                                                                                        No actual price increase however.

                                                                                    • alecsm

                                                                                      today at 5:30 PM

                                                                                      Right below the pricing it is stated that they plan to increase the prices in the near future.

                                                                                        • nchmy

                                                                                          today at 5:40 PM

                                                                                          "near future" is not "today"

                                                                                  • minraws

                                                                                    today at 5:11 PM

                                                                                    isn't it the same old pricing? did they increase V4 Pro pricing already?

                                                                        • indigodaddy

                                                                          today at 4:38 PM

                                                                          @dang - Pls merge this with https://news.ycombinator.com/item?id=49274018

                                                                          • book_mike

                                                                            today at 5:12 PM

                                                                            What I care about is whether the model is capable of the tasks I give it at the lowest cost. Right now I'm using Kimi-K3/GLM-5.2/Minimax. Sonnet is great but I burn through the tokens too fast. Opus 5 set to max is amazing and more intelligent than all of us. .998 of the time I don't need that kind of intelligence. I just need the job done.

                                                                              • okamiueru

                                                                                today at 5:17 PM

                                                                                How do you define intelligence? I encounter that kind of sentiment all too often, and I have to assume we go by wildly different understanding of what that might entail.

                                                                                  • bikemike026

                                                                                    today at 5:54 PM

                                                                                    If you read Opus 5's output, it is beyond the comprehension of virtually all engineers and developers. That is what I mean by intelligence. Math, science, and engineering are all contained in one model. We may be experts in one field. The model is an expert in everything that humans know.

                                                                            • nullbyte

                                                                              today at 5:46 PM

                                                                              Even though cost-per-token is low, Deepseek v4 tends to burn an immense number of tokens to accomplish tasks.

                                                                              • Readerium

                                                                                today at 5:23 PM

                                                                                V4 Pro has vision correct?

                                                                                  • trollbridge

                                                                                    today at 5:23 PM

                                                                                    No.

                                                                                • today at 4:19 PM

                                                                                  • nthypes

                                                                                    today at 5:32 PM

                                                                                    Still behind Kimi-K3 in almost half of the benchmarks

                                                                                    • LeonKnst

                                                                                      today at 4:41 PM

                                                                                      I find it interesting how much adoption seems to be influenced by momentum. Some of these Chinese models are surprisingly capable, but developers often default to the models that are already established as the “industry standard

                                                                                        • krlx

                                                                                          today at 5:43 PM

                                                                                          Well things may change soon. I've been testing Coding fulltime with Deepseek Flash this week to evaluate an eventual shift for the whole company away from anthropic. It has been quite positive and I can't wait to try pro tomorrow. If our data has to be used by either US or China, we might as well go the cheaper and unwalled garden. If only it supported image input ...

                                                                                          • sinuhe69

                                                                                            today at 5:11 PM

                                                                                            Well, one reason is that we always have to work with the quirks of each model. So, a know model is often preferred over a new/unknown one because we have to be vigilant again. (Negative) surprises are mentally exhausting in the long run. IMO, you can work much better when you know the model.

                                                                                            • spacebanana7

                                                                                              today at 5:09 PM

                                                                                              In an enterprise setting Chinese models are often discouraged due to political risk. They don't want to need to remove a model that's deeply embedded in their stack. And it's entirely feasible that the US gov bans federal contractors from using them in the next 6 months for example, or that EU AI safety rules effectively ban them too.

                                                                                                • trollbridge

                                                                                                  today at 5:38 PM

                                                                                                  Then run the DeepSeek or Qwen model on AWS GovCloud, etc., and you won't have any risk of exposure to "China".

                                                                                                  I'm not even sure what "EU AI safety rules" are. Can't people in the EU just use whatever they want?

                                                                                                  • BlackRabbit1

                                                                                                    today at 5:20 PM

                                                                                                    There are EU/US providers offering Deepseek/Qwen/Kimi/etc.-as-a-Service. With zero ties of their infrastructure to China.

                                                                                                    Fully compatible with the well known Antrophic API.

                                                                                                    You only have to replace the URL and your key.

                                                                                                      • odo1242

                                                                                                        today at 5:34 PM

                                                                                                        Based on what the political climate looks like nowadays it's entirely possible the US bans federal contractors from associating with any company that uses the models themselves, regardless of data provenance or where they are hosted. Or they create AI safety rules that make it impossible to release open source models (for example, making it so that closed-source models can be evaluated with a harness but open-source models need to pass the benchmark with the weights alone, which isn't really possible). Or they just declare Chinese models a security risk like TikTok (claiming that the model would be trained to respect Chinese interests).

                                                                                                        It may not be likely but it's definitely possible enough to be something people worry about.

                                                                                                • ianm218

                                                                                                  today at 5:19 PM

                                                                                                  I suspect if you follow dev groups in developing countries people are much more focused on token/ price efficiency.

                                                                                                  For funded startups it mostly just doesn’t matter a ton unless you are passing on inference in your product at scale

                                                                                                  • HawtAds

                                                                                                    today at 5:06 PM

                                                                                                    Hacker News is very Bay Area/US tech centric where spending a few hundred a month on AI is just pocket change. The weaker AI models with more questionable data retention policies are popular in developing countries. I think the new Facebook muse model will be similarly popular.

                                                                                                    • spacephysics

                                                                                                      today at 5:45 PM

                                                                                                      Most of my model usage comes from my work’s model selection (which is now down to just Claude models)

                                                                                                      I’ll try out the latest models, but mainly stick with Claude only because I’m most used to its quirks and how to work around them. I imagine this is part of these hyperscalers playbook.

                                                                                                      I will say though, I miss Sol model at work. It with Codex was amazing at first-shot understanding. Claude i need to scope out where to look otherwise a large portion of my token budget is eaten up

                                                                                                      • BlackRabbit1

                                                                                                        today at 5:06 PM

                                                                                                        A lot of it/infrastructure departments aren't aware that you can use Asian models hosted within the US or even EU.

                                                                                                    • yipinwong

                                                                                                      today at 5:31 PM

                                                                                                      Worse than Luna but more expensive than Luna. Sticking with Luna without sending my data to Deepseek (China)

                                                                                                        • Eueudhsbsj32

                                                                                                          today at 5:43 PM

                                                                                                          Unless you're Chinese, why would you care if they see your data?

                                                                                                          As an American, I'd much rather have my data kept outside the country than here where companies and the government have a lot more leverage over me.

                                                                                                            • today at 5:49 PM