\

DeepSeek-V4-Flash Update

648 points - today at 6:08 AM

Source
  • NitpickLawyer

    today at 6:25 AM

    This is more exciting than k3, IMO. Dsv4 models are extremely cheap to serve. Improving their capabilities has lots of downstream effects, as it becomes "good enough" for more and more tasks.

    DS was serving the pro version at extremely low prices for a long time, and they've had integrations with opencode & other providers, so they likely gathered a lot of data from real developers doing real tasks (on openrouter they were labeled as such). Now they can use those live scenarios to further post-train their models and improve them further.

    Can't wait to see if distilling k3 into dsv4 brings additional improvements. Anyway, having fast cheap models getting better is great for the community. Especially since these don't "go away" on a provider's whim. Whatever capabilities they get, can be used "forever" going forward. And, at least flash can be ran "at home" with <10k in hardware, which isn't really possible / feasible with glm/k3 larger models.

      • dotancohen

        today at 5:33 PM

          > Whatever capabilities they get, can be used "forever" going forward.
        
        "Forever" gets the scare quotes because it is implied only up until the Butlerian Jihad?

          • sigzero

            today at 6:54 PM

            I upvoted you just for the Dune reference.

              • idiotsecant

                today at 7:25 PM

                I feel like once they make a mega successful movie of a book we can stop doing the secret handshake.

                  • trvz

                    today at 8:12 PM

                    Especially when the movie is over forty years old.

        • chorizo

          today at 3:07 PM

          Starting to wonder if the free big pickle model on opencode has been DSV4F0731 for the past few months. It’s been incredibly fast and good.

            • kevincox

              today at 4:39 PM

              At least in the past it was GLM-4.6. IDK if it is ever changed.

              https://github.com/anomalyco/opencode/issues/4276

                • chorizo

                  today at 4:53 PM

                  I’ve heard that too, but since it’s a stealth model, I suspect they change it to whatever preview model provider that offers a free endpoint. In the past few months, I’ve gotten 4xx errors identifying the provider as DeepSeek.

          • dnhkng

            today at 6:41 AM

            Totally! This with DwarfStar delivers usable local AI (I hope!)

              • dannyw

                today at 9:14 AM

                Usable local AI has been here for a while, esp on say a 5090.

                You can’t treat Qwen3.6 like its fable, but if you prompt precisely and specifically it’s a great executor.

                I actually found it refreshing to use more of my brain for once, and actually have to think deeper about what I’m trying to do, and how to build it.

                • genxy

                  today at 5:12 PM

                  Parent is referring to https://github.com/antirez/ds4

                  • Tepix

                    today at 10:53 AM

                    What are your goalposts? Depending on your requirements, there have been many moments of usable local AI. More recent ones were gpt-oss 120b and Qwen 3.6 27b.

                • KaseyKim

                  today at 9:42 AM

                  hope that deepseek become better

                    • ilaksh

                      today at 10:52 AM

                      Have you tried the one that was just released?

              • f311a

                today at 7:46 AM

                I've been driving flash model for 90% of my tasks. It's better than pro (for unknown reasons), very cheap and fast.

                I try to keep changes under 1000 lines and drive architectural decisions myself, barely notice any difference compared to frontier models. The rest 10% is to spot bugs, security problems and to investigate better architecture, which flash can also do pretty well, I just cross check it.

                Faster iterations are way better for me, I hate waiting for 5-10 minutes on small changes. I tried to use recent versions of Kimi and GLM, but they use too much thinking for no reason and are pretty slow because of it. I also often feed a lot of data to it, without worrying about hitting the limits: dependencies (to find bottlenecks in them), logs, performance dumps and so on.

                Also, it will never complain about security guards, I've been using it to reverse engineer binaries.

                  • embedding-shape

                    today at 8:24 AM

                    > Also, it will never complain about security guards, I've been using it to reverse engineer binaries.

                    Maybe I'm using too weak language in my prompts, but none of the OpenAI models I've used via codex has refused to reverse engineer binaries, is it supposed to? I'm sitting right now reverse-engineering a 3rd party firmware together with Codex and haven't hit a single guardrail. Meanwhile, I see people complaining about it rejecting non-security related prompts, are things so individual on the platforms right now or what's going on?

                      • markasoftware

                        today at 8:42 AM

                        Have you completed the identity verification? It's much more lenient once you have

                          • embedding-shape

                            today at 8:44 AM

                            Oh yeah, back in the GPT3.5 days I think, that's probably it. Thanks for sharing your hunch :)

                            • flexagoon

                              today at 8:51 AM

                              Weird, I'm using 5.6 Sol through Chinese resellers and it reverse engineers stuff just fine

                                • johntarter

                                  today at 6:52 PM

                                  That's hilarious that Anthropic and OpenAI can't even secure their products from Chinese resellers, how are they supposed to secure their models from being distilled?

                                  • genxy

                                    today at 5:13 PM

                                    Because they have gone through identify verification, they are more likely to have done that to increase reputation score.

                                    • OkGoDoIt

                                      today at 5:46 PM

                                      I think I’m missing something, what do you mean Chinese resellers? As I understand it, it’s difficult to even access OpenAI in China, how could they be reselling it? Do you mean something like openrouter but a Chinese version or something?

                                        • markasoftware

                                          today at 8:38 PM

                                          They usually resell codex subscriptions as api so it's cheaper than the official api

                                          • Scoundreller

                                            today at 7:13 PM

                                            Proxies/vpns to enter/exit comms through non-banned countries.

                                        • dotancohen

                                          today at 5:42 PM

                                          Do you have any suggestions for resellers? My Gmail username is the same as my HN username if you're not comfortable posting that here.

                                          Thanks.

                                        • ngl999

                                          today at 9:40 AM

                                          Probably not real 5.6 Sol

                                            • flexagoon

                                              today at 11:10 AM

                                              How so? The code and reasoning patterns certainly match, so does the intelligence level. I probably wouldn't know if they're secretly serving Luna instead of Sol, but it's definitely an OpenAI model.

                                  • realusername

                                    today at 8:43 AM

                                    I got an account warning on OpenAI (waved after I complained) just because I was asking it how to root some >10 years old Android device.

                                      • dotancohen

                                        today at 5:44 PM

                                        Sonnet recently even suggested I root a three year old device and provided instructions. I suppose that is depends on use case. My use case was exporting data from an abandoned application running on an S24 Ultra.

                                        The idea was that I could continue to use the application in the future and export the new data, not that I would be able to recover the extent data already in there.

                                • mark_l_watson

                                  today at 11:28 AM

                                  This is my experience also: DeepSeek v4 flash is good enough for most of my work and I like the fast response times. I buy tokens mostly from FireWorks.ai in the US, but I also prepaid for a large chunk of tokens directoy with DeepSeek.

                                  I use OpenCode mostly (uses fewer tokens than Claude Code) and I am looking forward to the release of DeepSeek’s own coding harness.

                                  • regularfry

                                    today at 8:02 AM

                                    It's good but (at least on openrouter) it's got an annoyingly tight output token limit. So if it does get stuck in a reasoning pit, it won't work its way out of it in time.

                                    It's replaced the Kimi models for me though.

                                      • u8080

                                        today at 8:44 AM

                                        Try to use DS platform directly - cheaper and better than openrouter, no subscription

                                          • dghlsakjg

                                            today at 2:46 PM

                                            Openrouter has cheaper inference than deepseek for the prior version of v4 flash, which I suspect will happen within a day or two with this version, and they don’t offer subscriptions. Are you confusing it with opencode?

                                              • u8080

                                                today at 5:19 PM

                                                No, maybe I've checked too much time ago, but direct DS was cheaper for all versions - also I had issues with openrouter providers availability

                                                  • dghlsakjg

                                                    today at 6:09 PM

                                                    There are numerous providers offering full weight versions for a significant discount.

                                                    Availability and rate limiting can be an issue, though. I’ve found that constraining the providers works well to solve that though.

                                • lionkor

                                  today at 8:08 AM

                                  I use deepseek for a lot of my personal day-to-day agent needs, and I will simply put this here and let this speak for itself, last 30 days:

                                  - Cost: $4.55USD

                                  - API requests: 3,467

                                  - Tokens: 323,183,886

                                  And as an engineer who leads a small team, I have very high standards for quality, and these carry across to my personal projects where I use deepseek. It has not disappointed at all for coding or review tasks. For everything else, use another model.

                                    • troglodytetrain

                                      today at 6:52 PM

                                      DeepSeek is amazing, they are, from a cost/benefit literally an order of magnitude or more better than the 'SOTA' models, and yet no one really talks about them.

                                      I'm using them for my micro-saas, and they have made my niche economically profitable where as SOTA models are only slightly better for massively increased expense. Its truly impressive.

                                      Word of advice to anyone, not all your use of LLM tech needs to be code/dev work related.

                                      We are entering 'Web 4.0 era' or whatever you want to call it. Massive transformations of nearly every single business will and are being developed as the cost of intelligence as a commodity is falling through the floor...

                                      • a20eac1d

                                        today at 10:49 AM

                                        Can you give more info on how you use/prompt those LLMs for code review and what kind of prompts you use?

                                        I've had worse experiences doing it because the quality of answer has been quite bad, and I'm wondering if my methods are the reason.

                                          • lionkor

                                            today at 11:19 AM

                                            Yes, gladly! I have not yet open-sourced my skills etc., but I can give some insight and share a couple here.

                                            Review is a skill, as in, a SKILL.md with a folder full of references:

                                            - SKILL.md: https://gist.github.com/lionkor/161525be858d1d75db4c13c0f093...

                                            - references/output-contract.md: https://gist.github.com/lionkor/8c68e33becef7a21f8408c7dc119...

                                            - references/review-lenses.md: https://gist.github.com/lionkor/0a8b080fe45306213efddf3ebb75...

                                            - references/review-workflow.md: https://gist.github.com/lionkor/d2d374b133ceb7e3660bd530ee72...

                                            - references/section-rules.md: https://gist.github.com/lionkor/8a9e503adc7fd3697410cf021f27...

                                            I'm aware that almost all of this is prompt voodoo, and there's no guarantee for the review to find anything or everything, but making it a dedicated skill and thoroughly observing the output thinking, tool calls, and result, lets me adjust these over time and fill the weak spots with even more prompting.

                                            I use this skill by simply telling the agent something like "Review the changes on the current branch against origin/main, take special care with backwards-incompatible changes to the public API" or something like that.

                                            I use `pi` (pi.dev) with a subagents extension, so that I can ask the agent to invoke a subagent to do the review, on work that the agent did.

                                            For models, I use the highest possible reasoning on whatever model I feel like makes sense, usually this is GPT-5.5 or deepseek flash/pro, depending on the confidentiality of the codebase, on the highest reasoning always (for reviews).

                                            I've also had success with a review checklist, though it doesn't produce an easy to parse (for humans) output: https://gist.github.com/lionkor/054ac2cf241e0765eee2383f0dba...

                                            This is why my review skill mandates a very strict output contract. I need the output to be very easy to parse, and the output contract I've specified there does that.

                                            In general I let <whatever the latest model of OpenAI's ChatGPT is> author and review SKILL.md and similar large prompts, usually with a ruleset like this, which is a 1600 line research artifact from a long GPT 5.5 "Pro" research session on prompt engineering: https://gist.github.com/lionkor/71498794d0a7d72173fc58766f25...

                                            Does the review catch all issues? Not at all. Does it catch, usually more than one, important issue, across large changesets? Absolutely, and that's the point! :)

                                            Feel free to ask me any questions, I'm also happy to share more about my setup via email or add you or anyone else to my private repos with more of these.

                                              • rzerowan

                                                today at 5:27 PM

                                                Which languages are the reviews most successful on in your work so far? Im assumining js/python , would a similar output be feasible on lower level c++ or C#/Java.

                                                  • lionkor

                                                    today at 6:16 PM

                                                    No js or python, it's mostly C#, Rust, and C++. They're helpful/useful on all of those, I've also used the review skill on shell scripts, GitLab pipelines, Powershell scripts, C, and a couple other odd things. It's almost never useless to run it on something, it usually finds things, even if they're sometimes not very major.

                                                • kekebo

                                                  today at 2:57 PM

                                                  Thanks for sharing.

                                          • throwa356262

                                            today at 10:19 AM

                                            What harness are you using to achieve that level of token caching?

                                              • lionkor

                                                today at 11:19 AM

                                                I use pi, and, like the sibling comment, the caching ratio is fantastic. I work on C#, Rust, C, C++, shell scripting, and other areas.

                                                • k__

                                                  today at 11:11 AM

                                                  I'm using pi and my caching is ~99%.

                                              • brcmthrowaway

                                                today at 4:40 PM

                                                OpenRouter?

                                                  • lionkor

                                                    today at 6:19 PM

                                                    No, platform.deepseek.com for me! The caching is super important, not sure how open router performs, I haven't tried it.

                                                      • marcus_cemes

                                                        today at 7:17 PM

                                                        Works out of the box with OpenRouter for most models. Some providers are a bit flaky, but DeekSeek (provider) has been one of the most reliable for me, no problems hitting >99% CH.

                                                • dakolli

                                                  today at 9:18 AM

                                                  [flagged]

                                                    • lionkor

                                                      today at 11:27 AM

                                                      Aside from what squidbeak already said, I have to say the quality varies a lot, and goes above the "good" threshold enough times that using these models is worth the time and effort. Compared to something like GPT-5.4 to GPT-5.6 Codex models, they cost 50-100x less (fifty to one hundred TIMES less), and can do most annoying or well-specified tasks just as well.

                                                      The main difference is that deepseek is bad at prose, and Codex models are much more eager to use tools provided to them (which is usually fine, they often make like 3 todos via tool calls for a simple one step task which is silly though), and deepseek in general benefits from good instructions more than GPT models maybe do.

                                                      Deepseek becomes much better if you give it tools for asking clarifying questions, doing self-review with subagents, and so on. The more tools it uses, the better the signal-to-noise ratio in its context, and the more consistent the output.

                                                      You must still review 100% of the code as if it's trying to sell you insurance.

                                                        • mcbuilder

                                                          today at 2:26 PM

                                                          I love deepseek's (Pro) writing. I feel it's more nuanced and natural.

                                                          • ai_fry_ur_brain

                                                            today at 11:30 AM

                                                            [dead]

                                                        • squidbeak

                                                          today at 10:11 AM

                                                          An experienced software engineer (read his profile) praises the value he's found in Deepseek, and gives some real data showing how affordable that value is.

                                                          Then you, dakolli - out of generosity and minute-to-minute devotion to enlightenment - sacrifice time from your busy day to sit down (though perhaps that's been painful lately?) or stand up with your phone - and offer a profound, deeply thought-out counterpoint in the following form (and I'll paraphrase):

                                                          "Nah mate, it's shit. All LLMs are shit."

                                                            • CamperBob2

                                                              today at 5:32 PM

                                                              Look at his history -- he's just here to shitpost. Flag and move on.

                                                              • applicative

                                                                today at 12:05 PM

                                                                How one is affected psychologically by LLMs 'hallucinating' presumably closely tracks how one is affected by pathological liars. Most of the things pathological liars say is true and if you can put up with it, one can work with them and learn from them. But some respond with extreme hatred and avoidance. This seems like one of the responses that is reasonable ... as is acceptance and instrumentalization of the individual.

                                                                • _superposition_

                                                                  today at 11:15 AM

                                                                  Among countless other experienced engineers... The cognitive dissonance here is frightening.

                                                                  • dakolli

                                                                    today at 11:42 AM

                                                                    [flagged]

                                                                      • shock

                                                                        today at 12:36 PM

                                                                        > I really can't begin to describe how stupid I think you are for thinking that having that statement in your bio makes you a qualified engineer.

                                                                        Well, I guess having "Two PhDs, Three masters degrees. Expert in everything.." in your bio makes you very smart!

                                                                        • lionkor

                                                                          today at 12:25 PM

                                                                          I agree a bio is a bad proxy, and I'm not saying I'm amazing but I'm not THAT incompetent. Most of my GitHub is pre-LLM era; github.com/lionkor

                                                                          Edit: And like a lot of tools, HOW you use them is just about as important as the quality of the tool itself. Of course the tools can produce massive amounts of bad quality slop, they can also produce fast, focused edits that make sense.

                                                                            • dakolli

                                                                              today at 1:02 PM

                                                                              I'm not even trying to call you out, I'm sure you're solid.

                                                                                • lionkor

                                                                                  today at 1:10 PM

                                                                                  Oh I know, I'm just trying to say that I agree fully that a bio is a terrible proxy, while also showing what I think is a good proxy for skill (open source work before LLMs became this "good", for example). I'm a big proponent of show, don't tell, and LLMs have kind of ruined this.

                                                                  • ncphillips

                                                                    today at 9:26 AM

                                                                    For doing what?

                                                                      • dakolli

                                                                        today at 9:58 AM

                                                                        everything.

                                                                          • seanmcdirmid

                                                                            today at 10:10 AM

                                                                            Maybe it’s just your ability to use the tool and other people more proficient in using the tool are actually getting reasonable results from it.

                                                                              • ai_fry_ur_brain

                                                                                today at 11:28 AM

                                                                                [dead]

                                                                            • hlynurd

                                                                              today at 10:06 AM

                                                                              There's a small relief in knowing that these junk comments are at least not AI generated (in all likelihood.)

                                                                  • innis226

                                                                    today at 9:18 AM

                                                                    But doesn't it hallucinate a lot? Does that affect your workflow?

                                                                      • lionkor

                                                                        today at 11:22 AM

                                                                        It hallucinates plenty, about the same as Codex models and all other LLMs! I review all code it writes, thoroughly, check the test coverage, write tests myself, have other models/chats cross-check the work with a review skill (which I've shared in another comment in this thread).

                                                                        Usually unreviewed code only gets committed if I really don't care, like for one-off scripts, which I sandbox with github.com/lionkor/sbh or run as an unprivileged user.

                                                                    • maweaver

                                                                      today at 4:50 PM

                                                                      You are using DeepSeek's services directly? Doesn't that end up sending at least snippets/chunks of code to a server where it is subject to Chinese government data access laws? Even if I was okay with that, my organization would never be. And even if they were, our partners/vendors/customers would not be. I think that's the sticking point for a lot of people.

                                                                        • lionkor

                                                                          today at 6:20 PM

                                                                          Yes all of it goes straight to China! All of it is also open source or source available, or will be in the future. The stuff that isn't is trivial enough.

                                                                          For any work with protected intellectual property, I use other providers, for the contractual guarantees, but I think it would be silly to think that OpenAI or Anthropic are not training on literally all data they get. How could you ever tell if they did? They can just claim the data was mislabelled, or ignore the accusations. If you have serious IP, use only local models.

                                                                          • anigbrowl

                                                                            today at 6:06 PM

                                                                            So go with some other provider who hosts the models like OpenRouter if are worried about this.

                                                                    • kmarc

                                                                      today at 8:16 AM

                                                                      Essentially I'm running everything on flash now inside pi. With the correct set of MCP servers, context reducer tooling and skills it can implement any task I throw at it. Some sessions take 30+ turns, but it's fast and cheap; all this in an hour, with ~$0.5 cost.

                                                                      (TBH though, in my multi-subagent workflow I do use other, more expensive models for planning, reviewing, oracle-ing)

                                                                      I haven't used our slow opus subscription for weeks.

                                                                      (Also set up an OpenWebUi self-hosted chat that works from my phone, has some mcp and skills. fully replaced perplexity. Monthly cost ~$18 for hosting and subscriptions)

                                                                        • lionkor

                                                                          today at 8:35 AM

                                                                          I want to second this, I use the same setup (pi + deepseek, with lots of custom tools for tracking TODOs, doing things with less tokens, etc, and with a SOTA model for very difficult tasks) and it's all I need it to be.

                                                                          • peperunas

                                                                            today at 8:21 AM

                                                                            Do you have any recommendations of such extensions for pi?

                                                                              • sdesol

                                                                                today at 5:41 PM

                                                                                This is self promotional but I am working on making pi extremely enterprise ready with:

                                                                                https://github.com/gitsense/pi-brains/tree/staging

                                                                                The README is being worked on but the three videos should give you a good sense of what it can do. Pi is also what makes what I will demo in

                                                                                https://github.com/gitsense/chat/tree/update-readme

                                                                                possible. Since Pi exposes so much, it is very easy to build advanced tooling around it to help easily grok hundreds of tool calls to help you understand what they agent knows and what it has tried.

                                                                                • kmarc

                                                                                  today at 8:30 AM

                                                                                  Recommendation? No. Just go with the passive-aggressive advice "let pi build it for you". :-)

                                                                                  To be more constructive, what I did (as an experiencd SWE but a complete noob to agentic coding): went to pi.dev's extension marketplace and looked into all the new shiny stuff. Subagents, mcps, context and memory optimizers, skills. Using the most popular ones (not necessarily the best ones)

                                                                                  It was like 15years ago learning the new mindset of vim (and spending a ton of time to customize it to my workflow). My understanding is that Claude and opencode doesn't give you this flexibility.

                                                                                  Learning all these stuff drove me to also set up openwebui, and it was such a successful private project that I implemented it at work (with jira/confluence/bazel query access) and management said "we need this by tomorrow".

                                                                                  I believe the models matter not that much anymore. The "harness" does. (unless you just want to vibe code. Thebn, throw crap at fable and call it a day)

                                                                                    • javier123454321

                                                                                      today at 1:08 PM

                                                                                      I found that opencode absolutely lets you do everything that pi does. With a few niceties in the UI on top.

                                                                                        • epolanski

                                                                                          today at 4:05 PM

                                                                                          Opencode doesn't even allow providing your own system prompt.

                                                                                      • peperunas

                                                                                        today at 8:38 AM

                                                                                        Got it :-)

                                                                                    • rurban

                                                                                      today at 8:37 AM

                                                                                      Happy with oh-my-pi (omp)

                                                                                        • peperunas

                                                                                          today at 8:38 AM

                                                                                          I looked into it but, at least for me, it goes against the entire philosophy of pi :-)

                                                                                            • WithinReason

                                                                                              today at 9:16 AM

                                                                                              the entire philosophy of pi is to be extensible

                                                                                                • jeremyjh

                                                                                                  today at 10:47 AM

                                                                                                  omp is not really pi. Its a very distant fork of it. And its philosophy is completely different, it includes everything you need except workflow skills. Personally, I think it is the best agent coding tool by far, and its what I settled on after trying 6+ different tools.

                                                                                      • sergiotapia

                                                                                        today at 3:30 PM

                                                                                        I've had a lot of fun with https://omp.sh/ - consider it like zsh -> ohmyzsh

                                                                                        Lots of sensible defaults and good tweaks/settings you just don't worry about.

                                                                                        • try-working

                                                                                          today at 8:28 AM

                                                                                          pi-role-model

                                                                                      • kzrdude

                                                                                        today at 9:38 AM

                                                                                        Do you notice any improvements with this update?

                                                                                          • kmarc

                                                                                            today at 10:27 AM

                                                                                            Not sure, haven't tried it today.

                                                                                            Last night it single handedly implemented a feature after a grilling session, and came back with the red-yellow-green risk assessment points that I mostly saw with anthropic models. I had to check if I'm using the right model, but it was DS4Flash.

                                                                                            So maybe I was using it already?

                                                                                              • the_lucifer

                                                                                                today at 11:42 AM

                                                                                                Possibly depending on who's routing it, since Openrouter seems to have added a separate slug for this, and they don't have an alias router for the DS models yet, so it shouldn't have swapped over to it in case you use them for routing.

                                                                                                If you are directly using it through Deepseek, they do mention that the flash slug has migrated over

                                                                                        • rdsubhas

                                                                                          today at 9:01 AM

                                                                                          How can one set topp and temperature in pi?

                                                                                            • embedding-shape

                                                                                              today at 9:10 AM

                                                                                              You build it as an extension or create your own fork/copy of pi and add it. This basically goes for most things in pi, except the most basic stuff. It's basically meant for people who think "I'll just add that myself" rather than expecting it to be there out of the box or finding other's solution to it, for better or worse.

                                                                                      • wkcheng

                                                                                        today at 7:30 AM

                                                                                        If the benchmarks are real and reflect actual use, then this is an insane model. This 300B model outperforms the previous DS4 Pro preview model (1.8T params), and it looks like it outperforms GPT 5.6 Luna too. And it's still cheaper than Luna, even with the price decrease.

                                                                                        Crazy.

                                                                                          • gr_norm

                                                                                            today at 5:44 PM

                                                                                            OpenAI must've known this was coming, hence the Luna price drop. This competition is amazing!

                                                                                            • cbg0

                                                                                              today at 9:34 AM

                                                                                              On DeepSWE Deepseek is 54.4% and Luna is 67%

                                                                                                • benjiro29

                                                                                                  today at 10:43 AM

                                                                                                  But on Terminal bench, its

                                                                                                  * DS4 Flash: 82.7

                                                                                                  * GPT 5.6 Luna: 75.7

                                                                                                  For reference, that puts it on the third spot behind GPT 5.5 and Fable 5. For some reason GPT 5.6 Sol is not showing in the leaderboard. If it did, then DS4 Flash was number four.

                                                                                                  The thing is, even if Luna is better in DeepSWE and has the 80% discount. DeepSeek is still cheaper.

                                                                                                  --------------

                                                                                                  DeepSeek V4 | Flash GPT-5.6 Luna (New)

                                                                                                  --------------

                                                                                                  Input (Cache Hit) $0.0028 $0.02

                                                                                                  Input (Cache Miss) $0.14 $0.20

                                                                                                  Output $0.28 $1.20

                                                                                                  --------------

                                                                                                  Both Luna and Flash are heavy on the reasoning > output. And the cache hitrate + prices also matter.

                                                                                                  Reality is, you can not go wrong with Luna or Flash at those prices. And remember, DeepSeek V4 Pro is still in the rafters, what is ironically closer to Luna's new price.

                                                                                                    • flashblaze

                                                                                                      today at 12:37 PM

                                                                                                      Agreed. The only thing missing is multimodal capability

                                                                                          • dannyw

                                                                                            today at 9:12 AM

                                                                                            Deepseek and moonshot are the only two providers I consent to training for.

                                                                                              • arizen

                                                                                                today at 9:47 AM

                                                                                                Why not also Qwen?

                                                                                                  • dannyw

                                                                                                    today at 10:20 AM

                                                                                                    Qwen/Alibaba have stopped doing open weights releases for a while. No grudge or anything, I'm certainly not going to look at a gift horse in the mouth, but both DeepSeek and Moonshot have been very consistent with open weights as well as sharing actually detailed research.

                                                                                                    In terms of open research, China has absolutely overtaken the US.

                                                                                                      • anon373839

                                                                                                        today at 10:32 AM

                                                                                                        Qwen has a pinned tweet stating that 3.8 will be released as open weights soon. I guess it remains to be seen, though, if they’ll do the smaller model sizes or only the big 2.4T one.

                                                                                                          • SXX

                                                                                                            today at 10:50 AM

                                                                                                            Most interesting part IMO is whatever Qwen open weight will include image / video capabilities since its where they are strongest.

                                                                                                              • mcbuilder

                                                                                                                today at 12:27 PM

                                                                                                                Also base models versions are very useful for researchers, since they allow for a cold start to post training.

                                                                                                • siva7

                                                                                                  today at 1:57 PM

                                                                                                  CCP will be happy! Go on and share all your data with them..

                                                                                                    • neya

                                                                                                      today at 3:59 PM

                                                                                                      Yeah, because, your data is best collected by the US labs, right?

                                                                                                      US or American, don't trust anyone, these are open weight models. Host them yourself if you feel strongly about privacy (as you should). I honestly don't know of any OSS open weight models from the US labs as good as Kimi K3 or Deepseek V4 though.

                                                                                                      • fdsjgfklsfd

                                                                                                        today at 3:30 PM

                                                                                                        Yes, they trained on my open source code and Wikipedia edits and Stack Overflow answers that I shared freely, so it's only fair that they release their models as open source and share back to the community that created them. I only allow my training data go to open source models.

                                                                                                          • gr_norm

                                                                                                            today at 5:48 PM

                                                                                                            Yeah, at least when I share my training data with labs releasing their models openly (Chinese or American or otherwise) it's nominally so an even better open model will land in my hands in the future.

                                                                                                        • chorizo

                                                                                                          today at 3:09 PM

                                                                                                          I’m developing open source tools, so happy to have all my context traces be public if it helps improve future models.

                                                                                                          • anigbrowl

                                                                                                            today at 6:08 PM

                                                                                                            I will! I pay their sales too and consider it quite reasonable.

                                                                                                            • genxy

                                                                                                              today at 5:17 PM

                                                                                                              At least you get based models back at some point. You supply data, they supply compute and open weights.

                                                                                                              • bel8

                                                                                                                today at 6:26 PM

                                                                                                                Yes I will tyvm.

                                                                                                                They release the models back for free.

                                                                                                            • applicative

                                                                                                              today at 11:56 AM

                                                                                                              Xi going to shut down open-weighting of them in a matter of months. No one seriously doubts this. They will be too powerful and they will be gone.

                                                                                                                • satvikpendem

                                                                                                                  today at 2:41 PM

                                                                                                                  He explicitly said he wants open weight models at his recent speech at an AI conference in Shanghai: https://news.ycombinator.com/item?id=48970449#48970784

                                                                                                                    • applicative

                                                                                                                      today at 6:25 PM

                                                                                                                      The Chinese equivalent of the expression 'open weights' appears nowhere in this talk. You believed the Western press. Chinese AIs are not typically open. The most important by far is Bytedance.

                                                                                                                        • Petersipoi

                                                                                                                          today at 8:06 PM

                                                                                                                          The second China shuts down open weights is the second they lose the AI race. Nobody is going to use Chinese models if they're closed, besides hobbyists who don't have anything they care about getting stolen.

                                                                                                                  • gravypod

                                                                                                                    today at 12:15 PM

                                                                                                                    What would the benefit of this be? If China stops open weights, US labs still have the intelligence frontier. Maybe once Chinese models have speed, cost, and intelligence beat but right now they don't.

                                                                                                                    Undermining out entire economy by subsidizing the release of DIY versions of our main economic drive sounds like a huge win for China.

                                                                                                                      • today at 1:33 PM

                                                                                                                        • applicative

                                                                                                                          today at 1:41 PM

                                                                                                                          > Maybe once Chinese models have speed, cost, and intelligence beat but right now they don't.

                                                                                                                          Yes this is why I referred to 'months' and was downvoted by people who can't distinguish their politics from reality. You are restating exactly the text you are criticizing.

                                                                                                                          This is the nature of mechanical parrotlike repetition of propaganda:

                                                                                                                          a) You can have 'frontier' models with closed weights. Your sentence is basically a contraction within itself, again, the weight of ideology.

                                                                                                                          b) No one actually knows whether or not the closed weights of Bytedance are beyond all existing frontiers. Again the weight of ideology blinds you to the fact that the real 800 lb gorilla of Chinese Ai is more closed than Anthropic.

                                                                                                                            • ricardobeat

                                                                                                                              today at 2:39 PM

                                                                                                                              The latest ByteDance Seed model is behind Opus 4.6 in performance. It’s not even in the conversation atm.

                                                                                                                              The parent made a totally coherent commment that China closing off their open weight models would cede the frontier (and majority of AI users worldwide) to the USA. Nothing you said counters that point.

                                                                                                                                • applicative

                                                                                                                                  today at 6:26 PM

                                                                                                                                  Bytedance models completely pervade business use across China. It is the OpenAI of China, whether you think its a good model or not.

                                                                                                                              • gravypod

                                                                                                                                today at 3:52 PM

                                                                                                                                If they did have a super powerful secret AI why would avoid:

                                                                                                                                1. Selling access to the US? 2. Ship more and better software? 3. Talk about it publicly?

                                                                                                                                If you had a secret AI better than anything currently available the mere mention of this would crater the US tech investment sentiment.

                                                                                                                                They could even host these Chinese models through AWS and sell at-cost inference on Bedrock and obliterate the US model companies.

                                                                                                                        • anon373839

                                                                                                                          today at 12:37 PM

                                                                                                                          Reuters was reporting that rumor. And then Xi made a public appearance at a conference in Shanghai where he said the opposite of that rumor.

                                                                                                                            • applicative

                                                                                                                              today at 1:34 PM

                                                                                                                              No in that very speech Xi stated explicitly what amounted to: 'of course when we get a Mythos, it will be a state secret'.

                                                                                                                              In fact the overwhelming weight of AI use in China, the chatgpt so to say, is Bytedance's AI which is absolutely closed and uniquely opaque.

                                                                                                                              The press treatment of these matters was no good and they are slowly walking it back, e.g. NYT yesterday finally actually read the speech.

                                                                                                                                • anon373839

                                                                                                                                  today at 2:08 PM

                                                                                                                                  Do you have a source for that? The NYT article does not quote Xi making that statement, and I can't find a source for that. In fact, Xi delivered veiled criticism of the security posturing the US has done:

                                                                                                                                  > We should put in place laws and regulations, technological monitoring, early warning and emergency response systems in order to strengthen the line of security, prevent abuses and malicious use, and ensure that AI is always under human control. In the meantime, we should jointly oppose overstretching the national security concept in the field of AI and placing one country’s security over that of others.

                                                                                                                                  There was a blog post by a state-linked broadcaster cited in the NYT piece:

                                                                                                                                  > “China supports openness, but this does not mean it advocates for the unconditional proliferation of all capabilities,” the blog said.

                                                                                                                                  The article also quoted an American journalist:

                                                                                                                                  > “If these models do reach those dangerous capabilities, they are not going to let it be a free-for-all in terms of releasing them," Mr. Sheehan said.

                                                                                                                                  But a distinction has to be drawn between real dangers (which many people believe LLMs have not actually shown, to date) versus "dangers" hyped up marketing purposes or domestic regulatory-capture motives. Presumably the Chinese government is less interested in the latter.

                                                                                                                                    • applicative

                                                                                                                                      today at 2:14 PM

                                                                                                                                      "Dangers" may indeed be hyped up for marketing purposes or domestic regulatory-capture motives. But as Xi stated, they will become real.

                                                                                                                                      It is just a question of being unaffected by the motives of the speakers, which is what adults learn to do.

                                                                                                                                        • brazukadev

                                                                                                                                          today at 2:37 PM

                                                                                                                                          replying the request for sources with "as xi stated" without bringing the sources again is quite misleading. It seems you have nothing to back what you are saying.

                                                                                                                                            • applicative

                                                                                                                                              today at 6:28 PM

                                                                                                                                              "google is your friend"

                                                                                                                                              The person I am replying to actually believes that Xi actually spoke of "open weights", for example. This is the AI information space we live in.

                                                                                                                                          • anon373839

                                                                                                                                            today at 2:24 PM

                                                                                                                                            > But as Xi stated, they will become real.

                                                                                                                                            But the speech doesn't say that? I'm looking at the full text. For example:

                                                                                                                                            > We should take seriously the various types of inherent and secondary risks that AI may trigger.

                                                                                                                                            The risk portion of the speech is focused more on application-level risks than model capabilities. Which is a much more sensible regulatory framing than what Silicon Valley has proposed.

                                                                                                                                              • applicative

                                                                                                                                                today at 6:35 PM

                                                                                                                                                The claim that the speech was somehow supporting open weights was pure Western imposition. China does not support open weights generally, and Xi makes plain that some systems of weights - perhaps still future - would have inherent risks, same as Amodei says of Mythos, whether or not he is bull_ing about it.

                                                                                                                            • GTP

                                                                                                                              today at 1:34 PM

                                                                                                                              We can only wait to see if this is true. Another possibility could be that it was an answer to USA's government considering a ban on chinese models. In this way, he fueled the discussion around the importance of open weight models.

                                                                                                                                • applicative

                                                                                                                                  today at 1:40 PM

                                                                                                                                  Xi actually said that models of the strength to pose security issues will be permanent state secrets -- in the same speech that the first western takes translated as actually using the words 'open weights' which of course nowhere appeared. Now: when will "models of the strength to pose security issues" appear? Maybe my inference 'months' is wrong, but they seem to be advancing rapidly.

                                                                                                                                  The administration blather about banning open weights is characteristically confused. Xi has already stated (what is obvious) that he will ban security-endangering weights and keep them a state secret.

                                                                                                                              • darkwater

                                                                                                                                today at 12:15 PM

                                                                                                                                What makes you state this?

                                                                                                                                  • riskd

                                                                                                                                    today at 12:42 PM

                                                                                                                                    Decades of xenophobic propaganda

                                                                                                                                      • applicative

                                                                                                                                        today at 1:37 PM

                                                                                                                                        No, reading Xi's speech. They aren't going to give their Mythos - presumably a few months away - to the Sinaloa cartel or Uighur hackers or the US military or etc etc

                                                                                                                                          • h8hawk

                                                                                                                                            today at 2:34 PM

                                                                                                                                            They have already released their models (GLM 5.2 and K3). This is why Amodei and the rest of the effective altruist lobbies are trying to pressure the US government to both ban open-source models and pressure China to do the same.

                                                                                                                                            I’m curious to see why you’re so eager to ban open-source LLMs while being okay with Anthropic’s control. Do you think they’ll rebel against you?

                                                                                                                                            • sudosysgen

                                                                                                                                              today at 5:26 PM

                                                                                                                                              Mythos is Fable without safeguards. Open weight models are already essentially at the Mythos level, and with extra post training a well resourced actor could deploy would almost certainly be significantly better than Mythos at hacking. Whatever their threshold is, it's much higher than that.

                                                                                                                                          • 1234letshaveatw

                                                                                                                                            today at 1:18 PM

                                                                                                                                            Wow! Amazing observation! Wait, do you think there could have been any pro Chinese propaganda? Perhaps even some taking place at the present time? I think it could be possible?

                                                                                                                                              • Zetaphor

                                                                                                                                                today at 2:27 PM

                                                                                                                                                I'm sure there's plenty, but if you put on your rational hat and think about the motivations of the Chinese government you'll realize this claim makes no sense, and there's plenty of evidence and good reason saying the opposite

                                                                                                                                                  • 1234letshaveatw

                                                                                                                                                    today at 2:56 PM

                                                                                                                                                    you are probably right! nobody does more domestic discourse shaping than China, but good reason says none of that would take place externally

                                                                                                                                        • baublet

                                                                                                                                          today at 12:27 PM

                                                                                                                                          China bad

                                                                                                                                            • applicative

                                                                                                                                              today at 1:37 PM

                                                                                                                                              I'm a) quoting Xi's speech and b) projecting rapid China development to Mythos level. Either you think China will never get to 'Mythos level' because you think China bad; or you think they will release the weights, because you think China bad. Its incredible the weight of ideology over facts in this discourse

                                                                                                                                                • satvikpendem

                                                                                                                                                  today at 2:43 PM

                                                                                                                                                  You haven't quoted anything actually. Provide sources to your claims.

                                                                                                                                      • codemk8

                                                                                                                                        today at 5:58 PM

                                                                                                                                        Me.

                                                                                                                                        • 1234letshaveatw

                                                                                                                                          today at 1:23 PM

                                                                                                                                          That isn't the Chinese way. They are much more focused on undermine and extinguish. Just look at the European car industry- on its way to being non-existent after the market was flooded, bye bye manufacturing base. Undercut the US AI providers and wait them out, they will go private after they have a stranglehold

                                                                                                                                            • bwfan123

                                                                                                                                              today at 3:29 PM

                                                                                                                                              > Undercut the US AI providers and wait them out

                                                                                                                                              You could have said the same of linux - that it extinguished proprietary OSs at least for server usecases.

                                                                                                                                                • 1234letshaveatw

                                                                                                                                                  today at 5:42 PM

                                                                                                                                                  true, but the motivations of OSS contributors are benign (overall)

                                                                                                                                              • applicative

                                                                                                                                                today at 1:36 PM

                                                                                                                                                Even Xi's speech was as paranoid as Dario Amodei about the possibility of Chinese AI achieving something at the level of say Mythos. If you seriously believe such a thing will be on Hugging Face I really don't know what to say.

                                                                                                                                                  • fdsjgfklsfd

                                                                                                                                                    today at 3:36 PM

                                                                                                                                                    DeepSeek-V4-Flash-0731 scores higher than Fable 5 on Terminal-Bench, and Fable 5 is Mythos, correct?

                                                                                                                                                    • criley2

                                                                                                                                                      today at 1:44 PM

                                                                                                                                                      There is already something on HuggingFace at the level of Mythos. It's called Kimi K3 and it's running laps around Opus5, Fable5, and Sol56 at cybersecurity.

                                                                                                                                                      It's so good that the US government is rushing to ban all Chinese models as fast as they can.

                                                                                                                                                        • applicative

                                                                                                                                                          today at 1:50 PM

                                                                                                                                                          Fine, you believe Anthropic was lying about Mythos. In fact no view of the matter is needed to formulate the proposition above. It is clear China will soon have a model with the powers imputed to Mythos but which you deny of actual Mythos. It is a trial to deal with this level of ideological blindness; 'ideology has no outside'. If someone says 'when they get Mythos...', declare the possibility of Mythos to be the real lie. But it is quite possible and will be coming in months.

                                                                                                                                                            • zozbot234

                                                                                                                                                              today at 2:17 PM

                                                                                                                                                              Reading between the lines, it's fairly clear that Mythos Preview only got its reported results in offensive cyber thanks to a highly specialized harness and humongous amounts of test-time compute. We also don't know what kind of post-training/fine-tuning it got on cyber, Anthropic only stated that it wasn't specifically trained for cyber-offense but that leaves a lot of scope for training on other related things. Overall, the Mythos story is a lot less interesting than most people might like to claim - but note that this doesn't require any actual lying.

                                                                                                                                                              The interest around K3 is a lot more defensible because we actually know what the architecture looks like, how much effort it takes to get it to run, and what kinds of results it gets on cyber evaluations. And no, it's nowhere near "dangerous" enough to where people might honestly want to ban it for real safety reasons. It does a good enough job at fixing cyber issues, but that's hardly a safety concern.

                                                                                                                                                              • potwinkle

                                                                                                                                                                today at 2:42 PM

                                                                                                                                                                [dead]

                                                                                                                                                • ReptileMan

                                                                                                                                                  today at 12:56 PM

                                                                                                                                                  It is true. And only Chinese Han people can continue work on what is already released. But they will be forbidden. /s

                                                                                                                                                  Sarcasm aside - if the community can't get their shit together to continue improving what is currently public, well - we don't deserve free stuff and open weights.

                                                                                                                                          • ggcr

                                                                                                                                            today at 7:07 AM

                                                                                                                                            Woah, a 200B model competing with GLM-5.2 and getting close to Opus 4.8. Quite impressive.

                                                                                                                                            If those numbers translate well to its general capabilities, with the great caching DeepSeek has, I feel like this model will get tons of usage.

                                                                                                                                              • lostmsu

                                                                                                                                                today at 10:58 AM

                                                                                                                                                Not just 200B model, it is only 160GiB.

                                                                                                                                                • ignoramous

                                                                                                                                                  today at 12:34 PM

                                                                                                                                                  MiniMax's another lab that's known for relatively smaller models (their latest, M3 is 295b) that punch way above its weight.

                                                                                                                                              • heyalexej

                                                                                                                                                today at 7:25 PM

                                                                                                                                                Long story, I have humongous zai GLM 5.2 token budget that I'm using in a similar fashion as many comments explain here. GPT 5.6 or Fable 5 for planning, GLM for implementing, researching, extracting and many other tasks I consider grunt work. Very happy with the performance, speed isn't all that good though. I'd be curious to hear from someone who works with both, DeepSeek and GLM side by side.

                                                                                                                                                • Goranek

                                                                                                                                                  today at 7:47 AM

                                                                                                                                                  Kimi K3 (instead of Opus) for expensive stuff, DSV4 Flash for tasks (instead of Sonnet)?

                                                                                                                                                  Does this make sense?

                                                                                                                                                    • baalimago

                                                                                                                                                      today at 7:57 AM

                                                                                                                                                      Looks like they will release an updated version of deepseek-v4-pro soon, which most likely will beat kimi k3 at both intelligence and cost (judging by the vast improvements to dsv4 flash)

                                                                                                                                                      • geek_at

                                                                                                                                                        today at 7:51 AM

                                                                                                                                                        it does! I'm always amazed when I use DSV4 flash for coding or server checks and after an hour of working with it my (pure api call) balance is about 30 cents

                                                                                                                                                        • try-working

                                                                                                                                                          today at 8:02 AM

                                                                                                                                                          [flagged]

                                                                                                                                                      • thirtygeo

                                                                                                                                                        today at 7:33 AM

                                                                                                                                                        For both US and China models - what standard security checks and QaQc are you all doing? We're running small gamuts to test for unsolicited jailbreaks (model jailbreaks you) and incorrect records (Fake Accuracy - as Easter Egg or common thread) meaning falsified logic or information cooked in by the developers, rather than the training data speaking for itself

                                                                                                                                                        • baalimago

                                                                                                                                                          today at 7:37 AM

                                                                                                                                                          Very promising. So it will both keep the speed and reduced price, yet exceed performance of the quite sufficient deepseek-v4-pro?

                                                                                                                                                          Should be extending the lead in intelligence/cost index, as deepseek-v4-flash already were the most price efficient model, which now becomes even better. Although, in the deepseek APIs, the cost is leaking all information about codebases to China.

                                                                                                                                                            • ilaksh

                                                                                                                                                              today at 11:03 AM

                                                                                                                                                              Supposedly better than GLM 5.2 according to at least one benchmark.

                                                                                                                                                          • sim04ful

                                                                                                                                                            today at 8:42 AM

                                                                                                                                                            Why didn't they increment the version as atleast a patch update

                                                                                                                                                              • bel8

                                                                                                                                                                today at 6:27 PM

                                                                                                                                                                So whoever is using DeepSeek V4 flash gets a massive upgrade without having to change any config.

                                                                                                                                                                • crvdgc

                                                                                                                                                                  today at 1:46 PM

                                                                                                                                                                  My guess would be the version number tracks the architecture change, not the weight change.

                                                                                                                                                                  • rubslopes

                                                                                                                                                                    today at 11:54 AM

                                                                                                                                                                    I was also so annoyed by this. It makes communication all around difficult. Why don't they simply call it 4.1 flash?

                                                                                                                                                                    • today at 3:56 PM

                                                                                                                                                                      • kzrdude

                                                                                                                                                                        today at 9:43 AM

                                                                                                                                                                        Yes, and even if they seem to have updated endpoints in place they should assign the service a version number.

                                                                                                                                                                        • markasoftware

                                                                                                                                                                          today at 8:42 AM

                                                                                                                                                                          They want to be the next Google it seems

                                                                                                                                                                      • yewenjie

                                                                                                                                                                        today at 10:29 AM

                                                                                                                                                                        What was preventing them from calling it v4.1-Flash to distinguish it better?

                                                                                                                                                                          • lionkor

                                                                                                                                                                            today at 12:31 PM

                                                                                                                                                                            Sounds like it'll replace v4-flash, v4.1 would be nice to keep both available. On the other hand, it's nice to just get an improvement on anything that asks for "deepseek-v4-flash" without having to change the model string.

                                                                                                                                                                              • krapht

                                                                                                                                                                                today at 12:38 PM

                                                                                                                                                                                I think that's backwards. Anything that changes the performance of a model deserves a minor version bump. A new model has to be qualified before being pushed to production; but we don't get the choice here, just cross your fingers there are no regressions at all on all possible tasks the model might be asked to do.

                                                                                                                                                                                  • xbmcuser

                                                                                                                                                                                    today at 8:21 PM

                                                                                                                                                                                    its an open source model if it is important enough that you have to worry about a new version breaking something then why the fuck was that not running on your own servers this is not an Anthropic or OpenAI closed model that you only have access through an api

                                                                                                                                                                                    • halJordan

                                                                                                                                                                                      today at 4:00 PM

                                                                                                                                                                                      That's on you for using a preview/beta. It was properly qualified when it was released. No one complains when ios goes from beta to release; even though it will be ios 27 beta -> ios 27 next month

                                                                                                                                                                                      • lionkor

                                                                                                                                                                                        today at 12:51 PM

                                                                                                                                                                                        I should have worded it differently. Having a model selector as `string` seems like the wrong choice. At the very least, I should be able to select 4.X, like how we select dependencies, so people who rely on specific behavior can pin to 4.0 and I can say version >= 4 if I don't care.

                                                                                                                                                                                • halJordan

                                                                                                                                                                                  today at 4:01 PM

                                                                                                                                                                                  It went from "ds v4 preview" to "ds v4". That's enough of a distinction.

                                                                                                                                                                              • wg0

                                                                                                                                                                                today at 6:01 PM

                                                                                                                                                                                Don't know about the bench marks but I am getting Opus 4.7 level performance at fraction of cost with DeepSeek V4 Flash set to high. It is a reliable workhorse.

                                                                                                                                                                                • namuol

                                                                                                                                                                                  today at 5:51 PM

                                                                                                                                                                                  Can someone please explain how these models aren’t just fine tuned for benchmarks? I’m not plugged in to this space much but it seems like such an obvious problem…

                                                                                                                                                                                    • sroerick

                                                                                                                                                                                      today at 5:54 PM

                                                                                                                                                                                      They definitely are - but also people are using them pretty extensively for work. So ultimately you can't really fake "is it good". But there's no real measurements of that when a model is released, so we are stuck with benchmarks.

                                                                                                                                                                                        • namuol

                                                                                                                                                                                          today at 8:18 PM

                                                                                                                                                                                          I will continue to ignore the benchmarks.

                                                                                                                                                                                  • throwa356262

                                                                                                                                                                                    today at 1:01 PM

                                                                                                                                                                                    Gentlemen, start your DGX Sparks

                                                                                                                                                                                    https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731

                                                                                                                                                                                    • troglodytetrain

                                                                                                                                                                                      today at 6:47 PM

                                                                                                                                                                                      This is very exciting, my own niche micro-saas has already been able to make heavy use of DeepSeek-V4-Flash for my use case, I am looking forward to seeing how performance improves.

                                                                                                                                                                                      • alecsm

                                                                                                                                                                                        today at 12:43 PM

                                                                                                                                                                                        I've been using DeepSeek Pro for a while and Flash only for certain dumb tasks where I only need the speed of a LLM and not big brains.

                                                                                                                                                                                        I find the newest OpenAI and Anthropic models to be way better for big tasks that require many decisions but I don't like that anyway because I lose track of what's being done.

                                                                                                                                                                                        Knowing what I want for every prompt makes DeepSeek Pro the best LLM for me. It allows me to work relatively fast at a very low price.

                                                                                                                                                                                        • sqemo

                                                                                                                                                                                          today at 9:52 AM

                                                                                                                                                                                          DeepSeek is great for tasks and software I already know well. Even if it gets something wrong, I can usually verify it myself. But when I'm working with a programming language I'm not familiar with, I prefer using Codex or Claude.

                                                                                                                                                                                            • k__

                                                                                                                                                                                              today at 10:21 AM

                                                                                                                                                                                              Yeah, it needs quite some hand holding.

                                                                                                                                                                                              I didn't do much agent coding and had a mix experience.

                                                                                                                                                                                              1. It would build something that was in the spirit of what I wanted, but unusable in practice.

                                                                                                                                                                                              2. It would build something quite useful, but only the public APIs were nice, the deeper code layers would get more and more convoluted.

                                                                                                                                                                                              3. It would built what I wanted and it would have okay-ish code.

                                                                                                                                                                                              However, for 3. I also had to add a custom AGENTS.md, many more code example, extra repos as subtrees, and review any code that had new concepts.

                                                                                                                                                                                              Much more work, but still much less than typing it all by hand.

                                                                                                                                                                                          • arjie

                                                                                                                                                                                            today at 7:19 AM

                                                                                                                                                                                            Oh my goodness what an update. I need these weights. It's an incredible model for the size. The improved tool calling etc. should be able to make my harness way simpler. This runs at mega-speed on prosumer hardware (2x RTX Pro 6000).

                                                                                                                                                                                              • amunozo

                                                                                                                                                                                                today at 8:13 PM

                                                                                                                                                                                                How many tok/s are we talking about?

                                                                                                                                                                                            • mordae

                                                                                                                                                                                              today at 8:50 AM

                                                                                                                                                                                              I was just using it when it landed. It started reasoning more extensively from nowhere and precision went up a lot. It also changed its prose style for the better. Looking forward to weights.

                                                                                                                                                                                              • nickandbro

                                                                                                                                                                                                today at 6:59 AM

                                                                                                                                                                                                I wouldn't doubt GPT 4.6 Luna being in the top left quadrant's center on the Cost per Intelligence Index is not concerning for Liang Wenfeng. You have to remember DeepSeek v4 flash even though a bit cheaper, does not have vision abilities, which is a big draw for agentic tasks.

                                                                                                                                                                                                I admire DeepSeek's openness, but even they have been raising prices after their discounts.

                                                                                                                                                                                                  • minraws

                                                                                                                                                                                                    today at 7:07 AM

                                                                                                                                                                                                    They haven't raised prices though the plan was to increase it with peak hour usage for V4 Pro GA release, they didn't do that, so no price increases there I believe.

                                                                                                                                                                                                    As for vision yeah it sucks but Luna is also 2x input and 1.5x output for 1M context...

                                                                                                                                                                                                    That's around 0.4 in/1.8 out

                                                                                                                                                                                                    DSv4 is wayyy cheaper.

                                                                                                                                                                                                    And it's open now you have Luna at home if you have a decent set of GPUs you can run this on 2Sparks or one very expensive Mac or just like 6-8 5090s..

                                                                                                                                                                                                      • minraws

                                                                                                                                                                                                        today at 7:12 AM

                                                                                                                                                                                                        I know most folks can't afford it I am working on making it viable to rent shared hosting the biggest issue is data leak and prompt injection attacks with shared hosting. (Since the server owner connects to your main system via the coding agent)

                                                                                                                                                                                                        I guess using a ZDR provider is good enough for now.

                                                                                                                                                                                                    • gpugreg

                                                                                                                                                                                                      today at 10:18 AM

                                                                                                                                                                                                      According to the leaked call transcript, DeepSeek is working on vision for V4. Not sure when it will land though.

                                                                                                                                                                                                      • dudisubekti

                                                                                                                                                                                                        today at 7:09 AM

                                                                                                                                                                                                        Gpt 5.6 Luna cache read is $0.02 per mtok

                                                                                                                                                                                                        V4 flash cache read is $0.0028 per mtok

                                                                                                                                                                                                        That's not "a bit cheaper", just saying

                                                                                                                                                                                                          • nickandbro

                                                                                                                                                                                                            today at 7:20 AM

                                                                                                                                                                                                            That's a good point. Yeah their caching input is insane.

                                                                                                                                                                                                              • ignoramous

                                                                                                                                                                                                                today at 12:37 PM

                                                                                                                                                                                                                DeepSeek v4 API rates are only matched by Xiaomi's MiMo v2.5 series. Quite similar in capabilities too (at least before today's update).

                                                                                                                                                                                                        • re-thc

                                                                                                                                                                                                          today at 7:40 AM

                                                                                                                                                                                                          > I wouldn't doubt GPT 4.6 Luna being in the top left quadrant's center on the Cost per Intelligence Index is not concerning for Liang Wenfeng.

                                                                                                                                                                                                          The leaked interview has him saying it doesn't matter... as much as open source doesn't matter. There's enough in it for everyone right now and they aren't after everything.

                                                                                                                                                                                                          Perspective: DeepSeek doesn't have enough infrastructure to serve their target customers already.

                                                                                                                                                                                                            • nickandbro

                                                                                                                                                                                                              today at 7:47 AM

                                                                                                                                                                                                              Can you post the link to the leaked interview? From what I understand he has been pretty tight lipped for a guy who has a larger stake worth more in his company than Dario does in Anthropic.

                                                                                                                                                                                                                • ifwinterco

                                                                                                                                                                                                                  today at 12:21 PM

                                                                                                                                                                                                                  The culture is different, if you tried pulling any of the shit Sam and Dario do on a daily basis in China you'd get your wings clipped pretty quickly, much smarter to keep your head down as much as possible.

                                                                                                                                                                                                                  And maybe that's not a bad thing tbh because look at the state of the US right now

                                                                                                                                                                                                                    • re-thc

                                                                                                                                                                                                                      today at 2:21 PM

                                                                                                                                                                                                                      > The culture is different

                                                                                                                                                                                                                      That's partly. I'd say the other part is it's an open secret they have "illegal" Nvidia GPUs and other "secrets". There's a common understanding not talk about these things because it'd hurt everyone collectively.

                                                                                                                                                                                                      • amelius

                                                                                                                                                                                                        today at 10:32 AM

                                                                                                                                                                                                        Note: if you are having success with a model, then please post what you are using it for. Writing HTML/CSS is very different from writing Rust/C++ or doing maths.

                                                                                                                                                                                                          • fdsjgfklsfd

                                                                                                                                                                                                            today at 3:50 PM

                                                                                                                                                                                                            I use DeepSeek V4 Flash (before this update) for most things:

                                                                                                                                                                                                            - OpenCode for codebase editing: python scientific computing and LLM projects

                                                                                                                                                                                                            - Open Interpreter Classic (python version) for Swiss army knife terminal replacement one-off task type stuff.

                                                                                                                                                                                                            • f311a

                                                                                                                                                                                                              today at 10:51 AM

                                                                                                                                                                                                              I do Python, Go and Rust. Go works the best, I would say Python is worse than Go.

                                                                                                                                                                                                              Rust works perfectly fine, but when I use Rust, I usually pay closer attention to performance, so I have to guide it a bit, to improve cache locality, use simd, avoid unnecessary allocations and so on. Terra has the same issues with Rust. If you always prompt models to achieve the best performance, the code is usually no the one that I want, they optimize unnecessary/cold parts or blindly optimize stuff where compiler takes care of the optimizations already.

                                                                                                                                                                                                              See my other comment for more information on how I work with it. In short, keep the changes under 1k lines, context under 120k (ask it to use subagents), drive the architecture yourself.

                                                                                                                                                                                                              • sparse-Matrix

                                                                                                                                                                                                                today at 10:39 AM

                                                                                                                                                                                                                I've been having better-than-most performance using qwen3.6 to write python/flask.

                                                                                                                                                                                                                I tried using it to generate some rust code yesterday, and it generated much code but only ever came within 1 error of a testable build. The 4th or 5th full rewrite is sitting in the buffer right now.

                                                                                                                                                                                                                I'm currently looking to up my game with Bottlecap AI's return of qwen3.6, 'thinking cap'.

                                                                                                                                                                                                                Alleged to be twice as fast and superior at coding over extended sessions (vs. 3.6).

                                                                                                                                                                                                                We'll soon see.

                                                                                                                                                                                                                  • ilaksh

                                                                                                                                                                                                                    today at 11:02 AM

                                                                                                                                                                                                                    You should specify which model size and quant because there are many of both.

                                                                                                                                                                                                                • dandaka

                                                                                                                                                                                                                  today at 12:31 PM

                                                                                                                                                                                                                  Text classification, structured data extraction, rewrite

                                                                                                                                                                                                                  • bel8

                                                                                                                                                                                                                    today at 5:34 PM

                                                                                                                                                                                                                    flash for code:

                                                                                                                                                                                                                    c#, TypeScript, PHP, SQL, CSS, HTML.

                                                                                                                                                                                                                    also features, tests, fixes, refactoring and planning.

                                                                                                                                                                                                                    it's super fast, smart and dirt cheap.

                                                                                                                                                                                                                • Reubend

                                                                                                                                                                                                                  today at 6:57 AM

                                                                                                                                                                                                                  They're always very understated in their update descriptions. This is actually a HUGE improvement in the model's capabilities rather than just a small tweak.

                                                                                                                                                                                                                  • f6v

                                                                                                                                                                                                                    today at 8:28 AM

                                                                                                                                                                                                                    I'm thinking of using ChatGPT for making plans and V4-Flash for execution. Does anyone have good advice on pairing different models?

                                                                                                                                                                                                                      • davidjade

                                                                                                                                                                                                                        today at 3:28 PM

                                                                                                                                                                                                                        I stumbled into this same workflow idea. I hadn't really used anything other that Chat GPT Sol (medium) and when I wanted to start using code agents, I just asked GPT how to start. This evolved into having GPT do the brainstorming and writing the initial prompt for a project feature/change to hand over to Codex working in Sol light mode. These prompts where often a couple pages long as they had a lot of the architecture details worked out already. Codex would then come up with a plan and I'd take that back to GPT to review. GPT would make some suggestions, which I'd throw back to the agent and off it went.

                                                                                                                                                                                                                        I got super great results working this way. Maybe there is a better way to integrate the two modes though. But my first real project was end-to-end 100% working correctly from the first run - about 10,000 lines (including tests) of greenfield code. I did have it work in manageable chunks that I could easily review - not one-shotting the whole thing.

                                                                                                                                                                                                                        Now I'm thinking about plugging Deepseek into Codex to be the coding model and see how it goes.

                                                                                                                                                                                                                        • embedding-shape

                                                                                                                                                                                                                          today at 8:34 AM

                                                                                                                                                                                                                          Pairing non-APIs with APIs tend to be a hassle, and risky, as usually that breaks the ToC. I guess easiest for you to test if it's worth using the OpenAI API, is to manually copy-paste responses between wherever you run V4-Flash and the ChatGPT UI. What I've done in the past is basically .zip up the entire project directory, ignoring files from .gitignore, then asking ChatGPT Pro to inspect that and come up with a plan, then you paste that to where you have V4-Flash. Basically how we did "vibe pair programming" before the TUI agent harnesses appeared in the ecosystem :)

                                                                                                                                                                                                                            • yuzuquat

                                                                                                                                                                                                                              today at 9:21 AM

                                                                                                                                                                                                                              reminds me back in the day when we used to do version control with google drive. you generally don't need to do this anymore i don't think. it should be fairly trivial to open the codex cli in the project directory, ask it to do some analysis and planning -> save to md. then opencode with deepseek to read that md and start implementing. consider that some cross-model subagent workflows already exist: i use claude/fable to plan and there are openai plugins to directly spin up codex subagents. this latter approach gives claude/fable more directly control to babysit the codex agents whereas the former is more of a clearcut handoff

                                                                                                                                                                                                                                • embedding-shape

                                                                                                                                                                                                                                  today at 9:50 AM

                                                                                                                                                                                                                                  That approach won't give you access to the "Pro Mode" nor "Deep Research" though, which if you're paying for a Pro subscription already, "Pro Mode" is probably the top-top when it comes to "extensively researched answers & plans" across the entire ecosystem. That said, I do ask it to produce a .md plan that later is used with agents :)

                                                                                                                                                                                                                              • f6v

                                                                                                                                                                                                                                today at 12:00 PM

                                                                                                                                                                                                                                [dead]

                                                                                                                                                                                                                            • cthulberg

                                                                                                                                                                                                                              today at 9:35 AM

                                                                                                                                                                                                                              https://github.com/obra/superpowers

                                                                                                                                                                                                                              I use Opus/Sol with for /brainstorming, deepseek (on Pi) for /subagent-driven-development

                                                                                                                                                                                                                              I love it, the docs are easily editable and when I'm ready Deepseek is faster/cheaper then everything on my coding plans.

                                                                                                                                                                                                                                • leobg

                                                                                                                                                                                                                                  today at 10:19 AM

                                                                                                                                                                                                                                  Source reads like Tony Robbins for LLMs:

                                                                                                                                                                                                                                    Excuse
                                                                                                                                                                                                                                    "Too simple to test"
                                                                                                                                                                                                                                    
                                                                                                                                                                                                                                    Reality
                                                                                                                                                                                                                                    Simple code breaks. Test takes 30 seconds.

                                                                                                                                                                                                                          • ilaksh

                                                                                                                                                                                                                            today at 11:14 AM

                                                                                                                                                                                                                            I wonder when the antirez/ds4 group will have an update to their high accuracy 2 bit quant.

                                                                                                                                                                                                                            Although it's funny that I am thinking about that at all because I have a 2060 :P . My local inference is playing with Gemma 4 E2B and MiniCPM 5 1B.

                                                                                                                                                                                                                              • jtbaker

                                                                                                                                                                                                                                today at 7:28 PM

                                                                                                                                                                                                                                Looks like a couple of people are already on it: https://github.com/antirez/ds4/issues/635

                                                                                                                                                                                                                                • throwdbaaway

                                                                                                                                                                                                                                  today at 11:49 AM

                                                                                                                                                                                                                                  Objectively speaking, the 2 bit quant from antirez has very low accuracy. Meanwhile, his 4 bit quant does have decent accuracy, but is a bit pointless by being bigger than the full precision MXFP4 quant. Anyway, they all work fine in practice.

                                                                                                                                                                                                                              • Tepix

                                                                                                                                                                                                                                today at 10:55 AM

                                                                                                                                                                                                                                Sounds like a big improvement.

                                                                                                                                                                                                                                No mention of weights, just API. When will the updated weights be released?

                                                                                                                                                                                                                              • miyuru

                                                                                                                                                                                                                                today at 7:48 AM

                                                                                                                                                                                                                                Judging by the openrouter leaderboard ranking for today, it looks like Dv4F us more popular than mimov2.5.

                                                                                                                                                                                                                                https://openrouter.ai/rankings?view=day#leaderboard-table

                                                                                                                                                                                                                                These days cost per task is more important, and SOTA models have become expensive.

                                                                                                                                                                                                                                  • benjiro29

                                                                                                                                                                                                                                    today at 10:49 AM

                                                                                                                                                                                                                                    MiMo, the overlooked sidekick to the hero. Will be interesting to see what Xiaomi bring to the table.

                                                                                                                                                                                                                                    These massive jumps in cheap models, is really great times!

                                                                                                                                                                                                                                • KronisLV

                                                                                                                                                                                                                                  today at 7:13 AM

                                                                                                                                                                                                                                  Wonder how good the proper version of V4 Pro will be.

                                                                                                                                                                                                                                  I'm still considering pulling the trigger on the annual subscription of Kimi for K3 but it's sometimes slower than I'd like (at least when compared to Anthropic) even on their Vivace plan, and the token limits on the GLM Coding subscription for GLM 5.2 were too easy to hit.

                                                                                                                                                                                                                                  • HyperL0gi

                                                                                                                                                                                                                                    today at 11:42 AM

                                                                                                                                                                                                                                    Is anyone using DSv4 for their agents that is not related to writing code? Curious about use cases specially for someone using gpt-5.4 mini for classification, categorization, etc

                                                                                                                                                                                                                                      • freakynit

                                                                                                                                                                                                                                        today at 12:39 PM

                                                                                                                                                                                                                                        I use it to conduct thorough online researches. Plug-in some online search MCP (like the one I use: https://jerrysniffs.online ), and the flash models dig through the internet for dirt cheap..

                                                                                                                                                                                                                                        Most of the times, the total cost, including search API's, is less than $0.05 for full deeply researched output, and the research is actually good.

                                                                                                                                                                                                                                        • Lalabadie

                                                                                                                                                                                                                                          today at 1:37 PM

                                                                                                                                                                                                                                          It's been good for one-off cases in my limited experience. I would describe its behaviour as Sonnet-shaped, if that makes sense to you. Good answers but it often decides to reason a lot about simple things before getting to an output.

                                                                                                                                                                                                                                          At the speed Flash has on most providers, it doesn't really turn into a latency concern.

                                                                                                                                                                                                                                          • markab21

                                                                                                                                                                                                                                            today at 1:38 PM

                                                                                                                                                                                                                                            We use it at a moderate scale, self-hosted on B300 hardware. It's great :D

                                                                                                                                                                                                                                            QA analysis of voice transcriptions. Napkin math: we operate at 2-5% of the cost of running on Equiv Frontier, though this changes near-weekly because pricing is so volatile.

                                                                                                                                                                                                                                            It took us about a month to get the inference configured to achieve these numbers. But if you can get your hands on a pair of B300 GPUs and the context works, it's untouchable for price/performance.

                                                                                                                                                                                                                                            (B200 would work, but you don't have the B300's memory, which lets you run it on 2xGPU instead of 4xGPU... with Dspark, it's like magic)

                                                                                                                                                                                                                                            On a side note, for tasks that don't require the intelligence of DS v4 flash, we're using Nemotron-3-super with incredible success. I'm shocked we're not seeing more adoption of this model, given how easy it is to fine-tune and how blisteringly fast the nvfp4 version is. (A single B200 GPU can produce an insane amount of throughput with Nemotron 3 Super.)

                                                                                                                                                                                                                                        • vladukha

                                                                                                                                                                                                                                          today at 8:24 AM

                                                                                                                                                                                                                                          Where do you guys get deepseek? I'm hearing a lot of good reviews and want to try it with my pi config. from the deeepseek themselves, openrouter, or anywhere else? does it make a difference? [edit]: whoa it is really fast. will take some time to evaluate quality thou

                                                                                                                                                                                                                                            • u8080

                                                                                                                                                                                                                                              today at 9:57 AM

                                                                                                                                                                                                                                              Directly here: https://platform.deepseek.com/ Easy top-up and pay as you go. Availability and speed are very good and it is the cheaper option.

                                                                                                                                                                                                                                              • psibi

                                                                                                                                                                                                                                                today at 8:31 AM

                                                                                                                                                                                                                                                I've been using DeepSeek directly. I've heard from colleagues that using it via OpenRouter is slower, but I'm not so sure about that.

                                                                                                                                                                                                                                                • ticoombs

                                                                                                                                                                                                                                                  today at 8:40 AM

                                                                                                                                                                                                                                                  > does it make a difference

                                                                                                                                                                                                                                                  Probably not.

                                                                                                                                                                                                                                                  But Opencode-Go is a great solution for those who don't want to pay DeepSeek directly (or can't due to reasons)

                                                                                                                                                                                                                                                  Selfish referral code: https://opencode.ai/go?ref=R1AJZT4VBX

                                                                                                                                                                                                                                                  • embedding-shape

                                                                                                                                                                                                                                                    today at 8:37 AM

                                                                                                                                                                                                                                                    For hosted APIs, it's a lot cheaper to use their own infra, caching seems a hell of a lot better there compared to OpenRouter, and indeed the tok/s seems higher. They also have peak/off-peak pricing, so if you can hold off with your request, you get a pretty big discount.

                                                                                                                                                                                                                                                    Otherwise, if you're trying to run it locally, even really low quantizations like DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix seem to actually not be so dumb compared to smaller models with same quantization, might be worth a try if you're sitting on a lot of RAM/VRAM yet not industry-scale amount :)

                                                                                                                                                                                                                                                    • chronogram

                                                                                                                                                                                                                                                      today at 10:31 AM

                                                                                                                                                                                                                                                      I use it directly: https://platform.deepseek.com/usage

                                                                                                                                                                                                                                                      3rd party providers on OpenRouter can be cheaper but it's already so cheap.

                                                                                                                                                                                                                                                      • kzrdude

                                                                                                                                                                                                                                                        today at 9:41 AM

                                                                                                                                                                                                                                                        Opencode-go gives you $60 worth of DS V4 api usage for $10 per month. Right now I think it's hard to exhaust that when using flash exclusively, and plain API use might even be cheaper! Anyway, for DS usage it's a good deal.

                                                                                                                                                                                                                                                          • gpugreg

                                                                                                                                                                                                                                                            today at 10:34 AM

                                                                                                                                                                                                                                                            To add to this, the $60 only applies to DeepSeek-V4-Flash and a few other models. For DeepSeek-V4-Pro, the amount is $15.

                                                                                                                                                                                                                                                            https://opencode.ai/docs/go/#usage-limits

                                                                                                                                                                                                                                                            Previously, OpenCode Go had higher API prices for some models, but now they lowered the API price and simultaneously reduced the allowance.

                                                                                                                                                                                                                                                              • kzrdude

                                                                                                                                                                                                                                                                today at 10:54 AM

                                                                                                                                                                                                                                                                Thanks. These things change day by day I guess, AI is just moving fast (and I'm on vacation).

                                                                                                                                                                                                                                                                GPT 5.6 Luna is a new model in Go since I last checked, for example.

                                                                                                                                                                                                                                                            • Lalabadie

                                                                                                                                                                                                                                                              today at 10:30 AM

                                                                                                                                                                                                                                                              Opencode also have a ZDR (zero data retention) deal with them – if I recall correctly, that's not something you can enable as an individual DeepSeek subscriber.

                                                                                                                                                                                                                                                                • anon373839

                                                                                                                                                                                                                                                                  today at 10:41 AM

                                                                                                                                                                                                                                                                  I really like OpenCode - BUT: there is a loophole the size of Portugal in that ZDR language. All they say is that their providers follow a ZDR policy. I haven’t found anything promising that OpenCode themselves don’t retain Go usage data. Something to be mindful of.

                                                                                                                                                                                                                                                                  • gpugreg

                                                                                                                                                                                                                                                                    today at 10:36 AM

                                                                                                                                                                                                                                                                    Unfortunately, all mentions of ZDR have silently been removed from the OpenCode Go page today.

                                                                                                                                                                                                                                                                      • kzrdude

                                                                                                                                                                                                                                                                        today at 10:59 AM

                                                                                                                                                                                                                                                                        Thanks, any update here is important. I use them because of good data policies..

                                                                                                                                                                                                                                                                        I still find this today:

                                                                                                                                                                                                                                                                        > The plan is designed primarily for international users and provides stable global access. Your data will not be used for model training.

                                                                                                                                                                                                                                                                          • gpugreg

                                                                                                                                                                                                                                                                            today at 2:05 PM

                                                                                                                                                                                                                                                                            dax (coauthor) recently tweeted https://xcancel.com/thdxr/status/2083178051052155182

                                                                                                                                                                                                                                                                            > because we added the new deepseek which we do not yet have a ZDR with we cannot blanket say we offer ZDR

                                                                                                                                                                                                                                                                            I wonder how the website can make the statement that data will not be used for training.

                                                                                                                                                                                                                                                                            • Tepix

                                                                                                                                                                                                                                                                              today at 1:25 PM

                                                                                                                                                                                                                                                                              Do they do other things with the data, like selling it?

                                                                                                                                                                                                                                                              • lucianmarin

                                                                                                                                                                                                                                                                today at 12:16 PM

                                                                                                                                                                                                                                                                OpenCode harness gets the most out of DS V4 Flash model. You can implement any coding task, fast and cheap.

                                                                                                                                                                                                                                                                • Gigachad

                                                                                                                                                                                                                                                                  today at 8:28 AM

                                                                                                                                                                                                                                                                  I used it through openrouter. Plugged in to the vs code copilot bring your own key thing.

                                                                                                                                                                                                                                                                  Played around for a few hours and used up 80 cents of tokens.

                                                                                                                                                                                                                                                                    • Lalabadie

                                                                                                                                                                                                                                                                      today at 10:35 AM

                                                                                                                                                                                                                                                                      Anecdotal data from my own tests: allow only one provider if you want good cache usage on OR. 80 cents is probably 4x the price you should have paid.

                                                                                                                                                                                                                                                                      Providers' cache hit stats are available to consult, and only 1-2 of them behave properly if I remember correctly, zero if you request providers that don't store and train on sessions.

                                                                                                                                                                                                                                                                  • darkest_ruby

                                                                                                                                                                                                                                                                    today at 9:56 AM

                                                                                                                                                                                                                                                                    Openrouter

                                                                                                                                                                                                                                                                • wolttam

                                                                                                                                                                                                                                                                  today at 7:00 AM

                                                                                                                                                                                                                                                                  Hooray! This model makes me very optimistic about the future of local inference. The CyberGym score stands out to me.

                                                                                                                                                                                                                                                                  • egeozcan

                                                                                                                                                                                                                                                                    today at 7:31 AM

                                                                                                                                                                                                                                                                    Every time I want to have fun coding something with natural language processing, I use deepseek flash. It's just incredible for the price. I have a fairly popular app with 400 users that uses DeepSeek in the background and it still didn't hit even 50 bucks of usage in a month.

                                                                                                                                                                                                                                                                      • indigodaddy

                                                                                                                                                                                                                                                                        today at 11:27 AM

                                                                                                                                                                                                                                                                        Mind sharing the name? Sounds interesting.

                                                                                                                                                                                                                                                                    • PhilippGille

                                                                                                                                                                                                                                                                      today at 7:37 AM

                                                                                                                                                                                                                                                                      The previous V4 version wasn't called “Preview” by most inference providers. For example, the OpenRouter model slug was `deepseek/deepseek-v4-flash`. So now there will be confusion when someone talks about V4 Flash or when someone offers V4 Flash inference.

                                                                                                                                                                                                                                                                      Why not call it V4.1?

                                                                                                                                                                                                                                                                        • petu

                                                                                                                                                                                                                                                                          today at 8:18 AM

                                                                                                                                                                                                                                                                          It probably would be called 'deepseek-v4-flash-0731' in API

                                                                                                                                                                                                                                                                          edit: nope, at least deepseek kept "deepseek-v4-flash" and just updated model underneath. I guess preview is no longer worth serving with that release and you'd have to look through inference provider docs to see if they've updated, yeah..

                                                                                                                                                                                                                                                                            • PhilippGille

                                                                                                                                                                                                                                                                              today at 10:05 AM

                                                                                                                                                                                                                                                                              That's what I mean. On DeepSeek it's now just `deepseek-v4-flash`, while OpenRouter calls it `deepseek/deepseek-v4-flash-0731`, so now when someone talks about DeepSeek V4 Flash, like in benchmarks, or other inference providers, which version do they actually mean?

                                                                                                                                                                                                                                                                              The `-0731` style suffix is worse compared to a proper version bump like V4.1.

                                                                                                                                                                                                                                                                                • Macuyiko

                                                                                                                                                                                                                                                                                  today at 10:42 AM

                                                                                                                                                                                                                                                                                  What I and my team have been doing in papers is just to refer to the OpenRouter slugs. It upsets reviewers because they will complain it "is not sufficiently clear to a wider audience" but I do agree it's the cleanest approach. Also goes to prove how much power OpenRouter has actually...

                                                                                                                                                                                                                                                                              • kzrdude

                                                                                                                                                                                                                                                                                today at 9:46 AM

                                                                                                                                                                                                                                                                                Does it still say that it's an anthropic model, when asked? I would guess new post-training has fixed that.

                                                                                                                                                                                                                                                                                  • petu

                                                                                                                                                                                                                                                                                    today at 9:50 AM

                                                                                                                                                                                                                                                                                    Does Claude still say it's Deepseek, when asked?

                                                                                                                                                                                                                                                                                    https://news.ycombinator.com/item?id=49082022#49087112

                                                                                                                                                                                                                                                                                    How is that important? Maybe it does, so what?

                                                                                                                                                                                                                                                                                      • kzrdude

                                                                                                                                                                                                                                                                                        today at 10:41 AM

                                                                                                                                                                                                                                                                                        Not important, just a curiosity. I would expect that kind of knowledge would be pretty well burned into the weights, but what do I know about LLMs. Gemma 4 can answer this question well.

                                                                                                                                                                                                                                                                            • bermudi

                                                                                                                                                                                                                                                                              today at 9:43 AM

                                                                                                                                                                                                                                                                              DeepSeek being DeepSeek. v3 and R1 went over the same and had multiple versions

                                                                                                                                                                                                                                                                              • try-working

                                                                                                                                                                                                                                                                                today at 7:39 AM

                                                                                                                                                                                                                                                                                DeepSeek themselves called it `deepseek/deepseek-v4-flash`. Pro is still like that.

                                                                                                                                                                                                                                                                                  • PhilippGille

                                                                                                                                                                                                                                                                                    today at 10:08 AM

                                                                                                                                                                                                                                                                                    Yes that's my point. The old and the new version are different in capabilities, but now when someone talks about DeepSeek V4 Flash (in benchmarks, on inference providers), you don't know which exact version it's about.

                                                                                                                                                                                                                                                                                    Some providers like OpenRouter now call it `deepseek-v4-flash-0731`, but even in places like here on HackerNews people say things like "Sonnet is better than DeepSeek" without specifying a version or a reasoning effort, certainly no one will mention that `-0731` suffix when talking about DeepSeek V4 Flash.

                                                                                                                                                                                                                                                                                      • benjiro29

                                                                                                                                                                                                                                                                                        today at 10:54 AM

                                                                                                                                                                                                                                                                                        I think it does not matter. Most people who use these types of models are into the IT world, and will know/be informed very fast that there is a difference. People will likely also use the 0731 behind it.

                                                                                                                                                                                                                                                                                        The only issue i see, is 3th party providers that have not yet updated. But that is going to be a short time periode. There is no reason to not update.

                                                                                                                                                                                                                                                                            • w2seraph

                                                                                                                                                                                                                                                                              today at 6:40 PM

                                                                                                                                                                                                                                                                              This made my day !

                                                                                                                                                                                                                                                                              • flysoft

                                                                                                                                                                                                                                                                                today at 7:33 AM

                                                                                                                                                                                                                                                                                Finally have a model with usable intelligence, at a reasonable price. Can't imagine what Pro GA would look like, considering pro preview has only 1.6t parameters.

                                                                                                                                                                                                                                                                                • throwaw12

                                                                                                                                                                                                                                                                                  today at 8:57 AM

                                                                                                                                                                                                                                                                                  how different is their harness from Pi coding agent harness, is it possible to make an extension for Pi which can implement deepseek harness?

                                                                                                                                                                                                                                                                                  • storywatch

                                                                                                                                                                                                                                                                                    today at 7:43 AM

                                                                                                                                                                                                                                                                                    How's their performance in English prose? We are currently searching for cost effective ways to keep story wikis up to date.

                                                                                                                                                                                                                                                                                      • amunozo

                                                                                                                                                                                                                                                                                        today at 8:16 AM

                                                                                                                                                                                                                                                                                        I am interested in this too, as I think it could be these models are overly optimized for coding. Let me know if you figure it out!

                                                                                                                                                                                                                                                                                    • nathaah3

                                                                                                                                                                                                                                                                                      today at 8:55 AM

                                                                                                                                                                                                                                                                                      DS v4 flash has been my goto model for tasks in work. its been unsurprisingly fast and cheap.

                                                                                                                                                                                                                                                                                      • k__

                                                                                                                                                                                                                                                                                        today at 10:03 AM

                                                                                                                                                                                                                                                                                        I'd take more throughput while everything else stays the same.

                                                                                                                                                                                                                                                                                        • kamikazechaser

                                                                                                                                                                                                                                                                                          today at 7:41 AM

                                                                                                                                                                                                                                                                                          The flash variant is on par with Sonnet 5 on DeepSWE (54%). Big, if true.

                                                                                                                                                                                                                                                                                          • znnajdla

                                                                                                                                                                                                                                                                                            today at 10:05 AM

                                                                                                                                                                                                                                                                                            The conspiracy theorist in me wants to think that the 80% drop in GPT 5.6 Luna prices today is correlated with this update from DeepSeek. Perhaps OpenAI has already hacked its competitors with it's Mythos-like models and is aware of what competitors are doing and is able to react in advance.

                                                                                                                                                                                                                                                                                              • nchmy

                                                                                                                                                                                                                                                                                                today at 11:46 AM

                                                                                                                                                                                                                                                                                                Seems to me that it's the reverse. Open ai dropped prices to compete with deepseek etc and then deepseek released this update to negate Luna.

                                                                                                                                                                                                                                                                                            • spwa4

                                                                                                                                                                                                                                                                                              today at 6:27 AM

                                                                                                                                                                                                                                                                                              In case people want to run it, it's DeepSeek-V4-Flash-284B-A13B. So it should just barely run on a single B300, and it's small enough that it'll barely run on an M5 Max too.

                                                                                                                                                                                                                                                                                                • wolttam

                                                                                                                                                                                                                                                                                                  today at 6:31 AM

                                                                                                                                                                                                                                                                                                  It runs really well on 2 DGX Sparks - 60t/s

                                                                                                                                                                                                                                                                                                    • Tepix

                                                                                                                                                                                                                                                                                                      today at 1:32 PM

                                                                                                                                                                                                                                                                                                      Yes, the Dual DGX Spark looks like the sweet spot for this model for now. Good preprocessing speed. Lots of context. Fast enough for 1-5 devs perhaps. Around 8200€ as of today (used to be 6000€).

                                                                                                                                                                                                                                                                                                      Dual Strix Halo is much slower and current Macs with 256GB are both slower and more expensive (Mac Studio M3 Ultra 256GB around 12000€).

                                                                                                                                                                                                                                                                                                      To get something faster than the two Sparks you'd need to spend more than $22000 for a server with 2x RTX Pro 6000 at $10000 each.

                                                                                                                                                                                                                                                                                                      Beyond that you could get 2x AMD MI350P.

                                                                                                                                                                                                                                                                                                  • arjie

                                                                                                                                                                                                                                                                                                    today at 7:38 AM

                                                                                                                                                                                                                                                                                                    Not yet, right? That's the old DS V4 preview release. We're still waiting for the weights to come out.

                                                                                                                                                                                                                                                                                                      • benjiro29

                                                                                                                                                                                                                                                                                                        today at 10:56 AM

                                                                                                                                                                                                                                                                                                        Probably the same. When the same base model is trained, the weight do not tend to change a lot. GLM 5.0 > 5.1 > 5.2 are the same base model, that just kept being trained. Weights hardly change as a result. Think in the like few percentage points size difference.

                                                                                                                                                                                                                                                                                                          • today at 1:06 PM

                                                                                                                                                                                                                                                                                                    • lukan

                                                                                                                                                                                                                                                                                                      today at 7:46 AM

                                                                                                                                                                                                                                                                                                      "it'll barely run on an M5 Max "

                                                                                                                                                                                                                                                                                                      The max version I could order now with 128 GB?

                                                                                                                                                                                                                                                                                                      If so, the price for local inference would be 12 000 € vs 500 000 € for a B300.

                                                                                                                                                                                                                                                                                                        • NitpickLawyer

                                                                                                                                                                                                                                                                                                          today at 7:51 AM

                                                                                                                                                                                                                                                                                                          There's also the 2x spark way, which should be ~8k eur? Someone down the thread reported ~60tps for 2x sparks. That's totally usable for local inference.

                                                                                                                                                                                                                                                                                                          You can also do 2x 6kPRO in a workstation, for ~20k.

                                                                                                                                                                                                                                                                                                            • matrik

                                                                                                                                                                                                                                                                                                              today at 8:24 AM

                                                                                                                                                                                                                                                                                                              For the same performance, one could even go about 50% cheaper with 16 channel ddr4 + a rtx3090 for prompt processing.

                                                                                                                                                                                                                                                                                                              But still, even for mid level projects API is orders of magnitude cheaper, since you don't need to set it up and maintain it.

                                                                                                                                                                                                                                                                                                                • gpugreg

                                                                                                                                                                                                                                                                                                                  today at 10:44 AM

                                                                                                                                                                                                                                                                                                                  The memory bandwidth of the 2x RTX Pro 6000 Blackwell setup will be 10x higher, which should have an equivalent effect on the generated tokens per second.

                                                                                                                                                                                                                                                                                                              • spwa4

                                                                                                                                                                                                                                                                                                                today at 8:27 AM

                                                                                                                                                                                                                                                                                                                Currently the 3bit (and 2 bit) quant on DGX spark (on one of them) and the M5 Max should just start. Right now. (I'm hoping to get an M5 Max delivered on monday, let's see if it happens this time. It's 2+ months since I ordered now)

                                                                                                                                                                                                                                                                                                                The 4 bit quant technically fits (there's a 127 GB version) but ... obviously that's not going to work. It is so close though, surely someone will a way to do it.

                                                                                                                                                                                                                                                                                                            • reverius42

                                                                                                                                                                                                                                                                                                              today at 8:24 AM

                                                                                                                                                                                                                                                                                                              I'm running a useful quantization of the previous version of Deepseek-V4-Flash -- quite well but with so much fan noise -- on a MacBook Pro M5 Max with 128 GB.

                                                                                                                                                                                                                                                                                                              • spwa4

                                                                                                                                                                                                                                                                                                                today at 8:29 AM

                                                                                                                                                                                                                                                                                                                500k is for the 8x B300 version. Which is the only one you can buy atm. But technically a B300 card is more like 60k, just impossible to get.

                                                                                                                                                                                                                                                                                                        • XCSme

                                                                                                                                                                                                                                                                                                          today at 11:51 AM

                                                                                                                                                                                                                                                                                                          Can't really use it now, without giving away your data:

                                                                                                                                                                                                                                                                                                          > Trains: this provider may use prompts for training and may retain prompt data.

                                                                                                                                                                                                                                                                                                        • ra

                                                                                                                                                                                                                                                                                                          today at 8:03 AM

                                                                                                                                                                                                                                                                                                          What's the best way to run this on a 64GB M2 Pro?

                                                                                                                                                                                                                                                                                                            • petu

                                                                                                                                                                                                                                                                                                              today at 9:03 AM

                                                                                                                                                                                                                                                                                                              Weights are yet to be released (maybe in 24H, Deepseek has track record of releasing same day).

                                                                                                                                                                                                                                                                                                              https://github.com/antirez/ds4 is often mentioned for DS4F on Mac, but 64GB is likely not enough to achieve reasonable speeds (official weights should be ~160GB).

                                                                                                                                                                                                                                                                                                              • Tepix

                                                                                                                                                                                                                                                                                                                today at 1:27 PM

                                                                                                                                                                                                                                                                                                                DS4Flash has 284B weights. 64GB? No go.

                                                                                                                                                                                                                                                                                                                  • zozbot234

                                                                                                                                                                                                                                                                                                                    today at 2:25 PM

                                                                                                                                                                                                                                                                                                                    The native weights are 4-bit for the sparse experts, and they quantize to ~80GB with limited degradation in real-world performance. That's a viable target for 64GB with SSD streaming, though it will be slower than keeping the whole thing in RAM (especially on a M2-class machine with its slower storage).

                                                                                                                                                                                                                                                                                                            • sparse-Matrix

                                                                                                                                                                                                                                                                                                              today at 10:35 AM

                                                                                                                                                                                                                                                                                                              This may come as a surprise to a lot of AI concerns, but I have -zero- interest in paying for a model.

                                                                                                                                                                                                                                                                                                                • kzrdude

                                                                                                                                                                                                                                                                                                                  today at 11:06 AM

                                                                                                                                                                                                                                                                                                                  Then DS V4 flash is pretty relevant, because it's pushing prices down.. as well as being self hostable with a large enough rig.

                                                                                                                                                                                                                                                                                                                  • Tepix

                                                                                                                                                                                                                                                                                                                    today at 11:04 AM

                                                                                                                                                                                                                                                                                                                    Why mention it? Just download the weights when they become available.

                                                                                                                                                                                                                                                                                                                • sreekanth850

                                                                                                                                                                                                                                                                                                                  today at 12:44 PM

                                                                                                                                                                                                                                                                                                                  how this compare to luna high with reduced pricing.

                                                                                                                                                                                                                                                                                                                  • mekky16

                                                                                                                                                                                                                                                                                                                    today at 10:07 AM

                                                                                                                                                                                                                                                                                                                    if they were anthropic they wouldve just released it as a new model

                                                                                                                                                                                                                                                                                                                    • today at 6:31 AM

                                                                                                                                                                                                                                                                                                                      • truth_seeker

                                                                                                                                                                                                                                                                                                                        today at 9:54 AM

                                                                                                                                                                                                                                                                                                                        The magic of post training with valuable dataset

                                                                                                                                                                                                                                                                                                                        • sourcecodeplz

                                                                                                                                                                                                                                                                                                                          today at 8:48 AM

                                                                                                                                                                                                                                                                                                                          i've made a comparison between this and GPT Luna (recent %80 price drop)

                                                                                                                                                                                                                                                                                                                          https://x.com/SourceCodeplz/status/2083099712760987746

                                                                                                                                                                                                                                                                                                                          i prefer GPT-5.6 Luna honestly

                                                                                                                                                                                                                                                                                                                          • dakolli

                                                                                                                                                                                                                                                                                                                            today at 10:00 AM

                                                                                                                                                                                                                                                                                                                            Chinese labs rushing to release models this week, because it's inevitable that Washington regulates Chinese models in the next 4 weeks. All the US AI leaders have been taking trips to Washington this week, what do you think they're there for..

                                                                                                                                                                                                                                                                                                                            • Tepix

                                                                                                                                                                                                                                                                                                                              today at 10:50 AM

                                                                                                                                                                                                                                                                                                                              [dead]

                                                                                                                                                                                                                                                                                                                              • tosh

                                                                                                                                                                                                                                                                                                                                today at 7:22 AM

                                                                                                                                                                                                                                                                                                                                [dead]

                                                                                                                                                                                                                                                                                                                                • dnhkng

                                                                                                                                                                                                                                                                                                                                  today at 6:16 AM

                                                                                                                                                                                                                                                                                                                                  DeepSeek V4 Flash (Preview → 2026-07-31)

                                                                                                                                                                                                                                                                                                                                  • Terminal Bench: 56.9 → 82.7 (+25.8)

                                                                                                                                                                                                                                                                                                                                  • Toolathlon: 51.8 → 70.3 (+18.5)

                                                                                                                                                                                                                                                                                                                                  Compared to GPT-5.6 Terra:

                                                                                                                                                                                                                                                                                                                                  • Terminal Bench: Flash 82.7 vs Terra 78.4

                                                                                                                                                                                                                                                                                                                                  • Toolathlon: Flash 70.3 vs Terra 53.1

                                                                                                                                                                                                                                                                                                                                  • DeepSWE: Flash 54.4 vs Terra 69.6

                                                                                                                                                                                                                                                                                                                                  • Agents' Last Exam: Flash 25.2 vs Terra 50.4

                                                                                                                                                                                                                                                                                                                                  Trading blows with Terra, which is pretty interesting. No clear winner on these benchmarks, and wildy differeing scores. Very interesting!

                                                                                                                                                                                                                                                                                                                                    • villish

                                                                                                                                                                                                                                                                                                                                      today at 7:11 AM

                                                                                                                                                                                                                                                                                                                                      > Terminal Bench: Flash 82.7 vs Terra 78.4

                                                                                                                                                                                                                                                                                                                                      Terra 87.4

                                                                                                                                                                                                                                                                                                                                      https://openai.com/index/gpt-5-6/

                                                                                                                                                                                                                                                                                                                                        • benjiro29

                                                                                                                                                                                                                                                                                                                                          today at 11:00 AM

                                                                                                                                                                                                                                                                                                                                          https://www.tbench.ai/leaderboard/terminal-bench/2.1

                                                                                                                                                                                                                                                                                                                                          > 78.4

                                                                                                                                                                                                                                                                                                                                          The real score is always the official benchmark.

                                                                                                                                                                                                                                                                                                                                          We need to see later if DS4 flash 0731 is going to maintain the score but we need to look at the official benchmarks.

                                                                                                                                                                                                                                                                                                                                          Already seen a PR for DeepSWE to update the benchmark with 0731, so we can verify claimed vs official.

                                                                                                                                                                                                                                                                                                                                            • villish

                                                                                                                                                                                                                                                                                                                                              today at 3:23 PM

                                                                                                                                                                                                                                                                                                                                              The comment I replied to was comparing numbers released by each lab, and they made a typo with that specific benchmark.

                                                                                                                                                                                                                                                                                                                                      • throwaw12

                                                                                                                                                                                                                                                                                                                                        today at 6:59 AM

                                                                                                                                                                                                                                                                                                                                        Open flash model is competing against OpenAI's 'Sonnet' model at the price of GPT 3, I am really excited about this release, hopefully it holds up in real work as well

                                                                                                                                                                                                                                                                                                                                          • bayesianbot

                                                                                                                                                                                                                                                                                                                                            today at 8:11 AM

                                                                                                                                                                                                                                                                                                                                            IIRC GPT 3 was priced at per 1k tokens, had to check, the biggest GPT 3 model from OpenAI was $0.06/1k, so $60 / 1M. gpt-3.5-turbo was the first model after ChatGPT and that was $2 / 1M. And no caching. So not really in the same ballpark

                                                                                                                                                                                                                                                                                                                                        • Iolaum

                                                                                                                                                                                                                                                                                                                                          today at 6:34 AM

                                                                                                                                                                                                                                                                                                                                          Since they did this with their own harness I m not sure it's apples to apples comparison.

                                                                                                                                                                                                                                                                                                                                            • NitpickLawyer

                                                                                                                                                                                                                                                                                                                                              today at 6:37 AM

                                                                                                                                                                                                                                                                                                                                              > not sure it's apples to apples comparison.

                                                                                                                                                                                                                                                                                                                                              They're literally comparing the previous version of the same model with the new one. It's based on the same architecture, same pre-trained model, just different post-training. It doesn't get more apples to apples than this.

                                                                                                                                                                                                                                                                                                                                                • dnhkng

                                                                                                                                                                                                                                                                                                                                                  today at 6:40 AM

                                                                                                                                                                                                                                                                                                                                                  I think the commenter means the Flash vs Terra benchmarks.

                                                                                                                                                                                                                                                                                                                                                    • NitpickLawyer

                                                                                                                                                                                                                                                                                                                                                      today at 8:36 AM

                                                                                                                                                                                                                                                                                                                                                      Ah, my bad. Yeah that makes sense. They do say "The official V4-Flash natively supports the Responses API format and is specifically adapted for Codex.", so at some point someone will make a "same harness" comparison.

                                                                                                                                                                                                                                                                                                                                              • dnhkng

                                                                                                                                                                                                                                                                                                                                                today at 6:39 AM

                                                                                                                                                                                                                                                                                                                                                It will be fair if they release the harness though. I think now the future will be paired model-harness releases, not just weight dumps.

                                                                                                                                                                                                                                                                                                                                                The performance changes are so big with the right harness that is makes sense to engineer the harness and fine-tune the model to one another from the start.

                                                                                                                                                                                                                                                                                                                                            • yms_hi

                                                                                                                                                                                                                                                                                                                                              today at 6:53 AM

                                                                                                                                                                                                                                                                                                                                              I think it's better than GPT Luna.

                                                                                                                                                                                                                                                                                                                                          • try-working

                                                                                                                                                                                                                                                                                                                                            today at 7:40 AM

                                                                                                                                                                                                                                                                                                                                            Let's see how the market reacts.

                                                                                                                                                                                                                                                                                                                                            • Havoc

                                                                                                                                                                                                                                                                                                                                              today at 8:40 AM

                                                                                                                                                                                                                                                                                                                                              >benchmark results far exceeding V4-Pro-Preview:

                                                                                                                                                                                                                                                                                                                                              Wow that's crazy

                                                                                                                                                                                                                                                                                                                                              Good times for those that don't need strict data protection

                                                                                                                                                                                                                                                                                                                                                • blackoil

                                                                                                                                                                                                                                                                                                                                                  today at 8:47 AM

                                                                                                                                                                                                                                                                                                                                                  These being open provide much better data protection.

                                                                                                                                                                                                                                                                                                                                                  • kzrdude

                                                                                                                                                                                                                                                                                                                                                    today at 9:49 AM

                                                                                                                                                                                                                                                                                                                                                    Open weights means that there is a free market on providing inference using this model. Some of them will offer good data protection. (Well, hopefully)