\

Qwen 3.8 27B

351 points - today at 3:00 PM

Source
  • hypfer

    today at 4:32 PM

    Since it might be helpful to some, here's my current commandline for llama.cpp running on an RTX 4090 with my monitor moved to the iGPU to free up all of its VRAM.

    llama-server -m Qwen3.8-27B-IQ4_NL.gguf --mmproj mmproj-BF16.gguf -c 170000 --parallel 1 -ngl -1 --cache-type-k q8_0 --cache-type-v q8_0 -b 1024 -ub 512 --flash-attn on --no-context-shift --no-mmproj-offload --spec-type draft-mtp --spec-draft-n-max 5 --spec-default --cache-type-k-draft q4_0 --cache-type-v-draft q4_0 --threads 24 --jinja --reasoning on -fit off

    Identical to the qwen3.6 config. With a prompt like "svg owl" (which can reuse quite a lot compared with creative writing or similar, so ngram-mod shines), I get about 70-80t/s like this, with a memory overclock of about 1.5GHz

      • Aurornis

        today at 5:03 PM

        > --cache-type-k q8_0 --cache-type-v q8_0

        In my tests, even Q8 quantization for the KV cache comes with notable drops in performance for longer tasks. It does provide more context length in limited RAM budgets, but the longer context tasks are where KV quantization starts to show problems. It’s basically unnoticeable for simple and short tasks.

        > --spec-draft-n-max 5

        5 is a lot of tokens to draft. Are you really seeing acceptance rates to support that? When I tested it, 2-3 was the peak. Anything more started reducing performance except on highly predictable short outputs.

          • hypfer

            today at 5:06 PM

            Yes to both.

            The thing is that I can either use the q8 context, or have not enough context window, so I just live with whatever degradation there is. The same can be said about the IQ4_NL. I would not go any lower though.

            As for the draft count, indeed that depends on what you do with it, but for coding, reverse engineering and that kind of stuff it does pay off in my testing, though 5 is really pushing it, but the 4090 has so much compute.

            Last logline I saw scroll by right now had 47% acceptance rate for 4th and 28% for 5th, but not sure how representative that is. I think when tuning 3.6, I saw more like 33%? But not 100% sure.

        • reilly3000

          today at 4:43 PM

          Thanks for posting! Have you had any success with running without kv cache quantization? Is there a noticeable difference in quality without any? I would assume that would eat into context but 170k is pretty generous!

            • hypfer

              today at 4:46 PM

              According to this shitty vibecoded thing "I" built https://hypfer.github.io/will-it-fit-llama-cpp/ (and I guess according to math too), FP16 K/V would give me something like 90k context at the same model quant, which doesn't really fit my usage.

              But maybe someone else has experience to share there

                • nubg

                  today at 5:06 PM

                  just to clarify. yes YOU built it. just because you used some tool doesn't mean the idea, prompting, reprompting, babysitting was not your creative input and effort.

                  put differently, if you put a random person infront of whatever model you used (say, a 50yo receptionist at a pharmacy in india), they would not have been able to create that, because they would have lacked the motivation, idea, background knowledge, taste, etc to create such a thing.

                    • bilekas

                      today at 5:16 PM

                      You sound like your trying to reassure yourself of something.

                      I sure hope my boss doesn't think he built my work! He'd probably get fired pretty quickly during on call!

                        • formerly_proven

                          today at 5:59 PM

                          > I sure hope my boss doesn't think he built my work!

                          Most managers do though?

                          • today at 5:18 PM

                        • williamcotton

                          today at 5:16 PM

                          I generally agree and expect this to be the case from a legal perspective.

                          Legal questions of authorship are going to have to be established in terms of doctrines like SSO [0] and AFC [1]. Currently the incredibly sparse caselaw around this has yet to involve such non-literal notions of copyright.

                          [0] https://en.wikipedia.org/wiki/Structure,_sequence_and_organi...

                          [1] https://en.wikipedia.org/wiki/Abstraction–filtration–compari...

                          • effdee

                            today at 5:32 PM

                            Some people will now argue it was the chisel—not Michelangelo—who created David.

                              • b112

                                today at 5:43 PM

                                No, it's the difference between management and direct work.

                                None would claim they chiseled anything, if it was 3D printed. They may claim they designed something.

                                • smallmancontrov

                                  today at 5:43 PM

                                  "Carve me a naked guy. Make no mistakes."

                              • tinfoilhatter

                                today at 5:24 PM

                                So if I hire an artist and am a motivated individual, have an idea for a painting, have background knowledge about paintings and have taste in paintings and can offer a critique of the painting as the artist paints it, then somehow I created the painting?

                                Absurd logic. The AI built the website.

                                  • sampullman

                                    today at 5:30 PM

                                    I think in that case it's fair to say you created the painting with the artist, even if the artist should get majority credit. I don't like the analogy though, to me it feels more like you're a project manager directing a team of genius but single minded interns.

                                    • mixologic

                                      today at 5:45 PM

                                      Nothing absurd about that. What do you think an "Executive producer" is? A "Director" ? Does Peter Jackson get credit for creating the Lord of the Rings Trilogy films? Christopher Nolan for his films? But did he make them ? No, it was the collective effort of thousands of individuals all working under their direction.

                                      Just like if somebody creates software today, and the end result is generated by the collective effort of thousands of agents, the "Director" still gets credit.

                                      • rob

                                        today at 5:35 PM

                                        I just read through a couple of your posts that weren't dead or buried, and it seems like you're pretty anti-AI. You should really start to have an open mind towards it. It's going to be the future (if it isn't already), and as you continue to get older, you're going to really wish you spent your time right now learning and embracing the technology instead of being so against it. A lot of the skills and things that you're holding on to right now might not be relevant by then, but you'll be at a disadvantage from not keeping up with the industry and need to play catch-up.

                                        • alienbaby

                                          today at 5:51 PM

                                          You know how many pieces of art Damien Hurst creates himself Vs his studio assistants creating them under his direction?

                                          For example, of his 1500 spot paintings, he only actually made 5 of them.

                                          It's not uncommon at all for artists to work this way.

                                          • williamcotton

                                            today at 5:53 PM

                                            There have been plenty of workshops where artists hire assistant painters while maintaining authorship over the works themselves, from Rembrandt to Warhol to Hirst.

                                            • nubg

                                              today at 5:47 PM

                                              no, because there is another human involved.

                                              llms are not human.

                              • D4Ha

                                today at 4:52 PM

                                Do you find it useful or worthwhile to split a large LLM across two GPUs on a desktop?

                                If you've tried it, what worked well and what didn't? I'm especially interested in mismatched VRAM setups, e.g. a 16 GB GPU + a 24 GB GPU.

                                How much overhead did you see from inter-GPU transfers, and did the extra usable VRAM outweigh the performance hit?

                                  • evanreichard

                                    today at 5:55 PM

                                    As opposed to loading it up in RAM + VRAM? Pretty much always better to split it up to multiple GPUs. My priorities are load up all available VRAM, then offload MoE experts to RAM (if possible), then offload other layers.

                                    I use an RTX 3090 (24GB) and a GTX 1080ti (11GB). Just the 3090 for 3.8 I get maybe 60tg/s (UD-Q4 quant), for both I get around 40tg/s (UD-Q6). Not apples to apples though considering it's different quants.

                                    Here's my config: https://gitea.va.reichard.io/evan/nix/src/branch/master/modu...

                                    • giyanani

                                      today at 5:37 PM

                                      It depends on what model you’re running, and for what workload. For personal use (one or two convos at a time) with models that fit in gpu memory, pcie bandwidth doesn't really matter. Just try and be on gen 3 x8 or higher.

                                      Llama is decent at auto optimizing it if you let it use both gpus. It’ll split the workload so the contiguous layers are all on one gpu. Once the model is loaded, you only transfer weights between gpus once (per token?), at the layer boundary.

                                      I run on an 8gb 3070 and 12 gb 3060, and the only weird thing is that the weaker card gets more layers (and therefore work) because it has more ram.

                                      Oh, if you’re barely fitting the models into your vram, you may need to explicitly adjust the layer balance between cards — sometimes it fails to realize it should have put certain things (like draft models) on the other card so you can fit one more layer in.

                                      • usagisushi

                                        today at 5:38 PM

                                        A hetero-GPU setup is definitely cost-effective if you don't strictly require the raw speed of a top-tier card like 5090. Just keep in mind that the total throughput will also be bottlenecked by the slower card.

                                        To provide some anecdotal data, here is how my 5090 + 3060 setup performs with Qwen 3.8 27B (Unsloth's UD-Q4 with MTP):

                                          Single 5090: 101 t/s (TG), 2650 t/s (PP)
                                          5090 + 3060: 53 t/s (TG), 1700 t/s (PP)
                                        
                                        For reference, here are also some numbers from my 4060ti + 3060 (16GB + 12GB) setup. [0]

                                        [0]: https://news.ycombinator.com/item?id=48700091

                                        • mechagodzilla

                                          today at 5:20 PM

                                          Yes. I can split a model like this across 3 GPUs (a 1080 with 8GB and two Titan Vs with 12GB), and it's much faster than running it on 36 CPU cores. As long as it fits in aggregate VRAM, it seems very advantageous to do so.

                                          • mips_avatar

                                            today at 5:18 PM

                                            Depends on your pcie connection. If they're both x16 then it's pretty low overhead, x8 is ok, but x4 is too slow. Also it's a bit tricky getting an optimal setups with mismatched vram, I think you could probably still make use of the full vram if you're clever but it's trickier.

                                              • today at 6:01 PM

                                                • nullc

                                                  today at 5:23 PM

                                                  for layer parallelism (e.g. to get more vram) the bandwidth between layers is essentially nothing (like 16kb per token I think), so I don't think x4 would even be a problem!

                                                    • ericd

                                                      today at 6:03 PM

                                                      Good point. It's much more of an issue when running dense models with tensor parallelism. In that case, I'd look for an MoE model instead.

                                              • bilekas

                                                today at 5:19 PM

                                                I haven't tried this either but I'm guessing if you could pool the GPU memory over whatever the kids are using these days, I think it was SLI back in my day. The GPU memory should still be faster than the RAM?

                                                • today at 4:54 PM

                                                  • CamperBob2

                                                    today at 5:00 PM

                                                    That's a very deep rabbit hole involving PCIe topology on both the hardware and software (NCCL) side, among other things. It's too system-specific to answer directly, but the entrance to said hole can be found at https://github.com/local-inference-lab/rtx6kpro/blob/master/... .

                                                    Disregard references to RTX 6000 cards, most of it is generally applicable to all multiple-GPU boxes.

                                                • bmitc

                                                  today at 4:53 PM

                                                  Lol at that command. Why is this stuff so hard to run locally? I've spent a few days trying to figure it all out and haven't been able to. LM Studio doesn't work behind proxies. Ollama is confusing and doesn't seem to support Qwen3? And Llama.cpp is your command.

                                                  I just want to run `<some-command> <model-name>` with some default parameters set and for it to run locally.

                                                    • hypfer

                                                      today at 4:55 PM

                                                      What makes you say that it would be hard to do that?

                                                      It's long, I guess, but not cryptic.

                                                      You tell llama server where the model is, which context size to use, what to use for the K/V cache quant, that it should do MTP, tune some MTP parameters, and that's kinda it.

                                                      Perfectly logical blocks with all the model-specific weirdness (that does exist!) abstracted away.

                                                      You could also just run -m <modelfile> and let llama-server do the right-ish thing. The defaults are probably fine, but not how you squeeze out these exact numbers. I think at least. I've never tried. My hubris stopped me from trying auto configs.

                                                        • porphyra

                                                          today at 5:02 PM

                                                          > llama server where the model is, which context size to use, what to use for the K/V cache quant, that it should do MTP, tune some MTP parameters, and that's kinda it.

                                                          isn't that the hard part? You know the ballpark ideal values for these many parameters since you're a knowledgeable expert but the vast majority of people are just like "I want AI" and have no idea what all the jargon even means.

                                                            • solenoid0937

                                                              today at 5:06 PM

                                                              Sure, but front ends like LMStudio exist for that crowd

                                                              Otherwise, if you're a programmer setting up a local harness, it only takes like 20-30 minutes to learn what the right parameters are.

                                                              It's very model, hardware, and use case dependent which is why a one size fits all solution doesn't work

                                                              • manquer

                                                                today at 5:07 PM

                                                                Why would they wish to handcraft this ? That is what agents are for ?

                                                                They could ask your current agent to a) search for this type of content online for the optimal setup for their hardware b) have the current agent/harness spin it up have it verify the config run few experiments.

                                                                Sure AI may make mistakes, or won't get the best possible config probably, but it certainly do a good enough setup, this is a task with feedback on whether the server crashed or poor performance easily measured so the agent can do a pretty good job.

                                                                  • hypfer

                                                                    today at 5:26 PM

                                                                    > Why would they wish to handcraft this ?

                                                                    Because this is kinda the one new thing that arrived in the technology scene, so getting at least some amount of understanding of its "inner" workings might prove useful in the future.

                                                                    Beside that, it is also just.. interesting? It's fun tuning the machine to see it improve. For some, anyway.

                                                                      • manquer

                                                                        today at 5:33 PM

                                                                        The people OP mentioned about "just want AI" .

                                                                        The pain point they raised is this is too complicated for people who just want to get started, that is not true anymore.

                                                                        It is certainly fun to fine-tune and setup if you like do something like that, however the need to do it hardly is a barrier for those who don't want complexity as OP imagines.

                                                                        Lower level API/interfaces should not be a barrier for people if they are apply framing that way. More and more people are thinking agent native so this is not really a issue.

                                                                          • hypfer

                                                                            today at 5:48 PM

                                                                            > More and more people are thinking agent native so this is not really a issue.

                                                                            Why this headache inducing lingo tho? What does that even mean, and why should I sign up for your webinar about that?

                                                                • bmitc

                                                                  today at 5:06 PM

                                                                  This is exactly it. I already have broad access to Claude, Gemini, GitHub Copilot. I want to use open models on automated tasks that chew up tokens but where I don't necessarily need the best-in class models and UX.

                                                                  For Claude, I setting a single config file and then download and run Claude Code CLI. Even easier for the GitHub Copilot CLI.

                                                          • Aurornis

                                                            today at 5:10 PM

                                                            Start by copying the command line from the Unsloth guides.

                                                            You don’t need to fine tune all of those parameters to get started.

                                                            It’s really easy to ask an LLM to adjust the command line if you can’t be bothered to read the help out. Copy the help output into the LLM and tell it your goal.

                                                            > Ollama is confusing and doesn't seem to support Qwen3?

                                                            Typing “Ollama qwen3” into Google takes you right to this page:

                                                            https://ollama.com/library/qwen3

                                                            If even Googling for basic Ollama support is too hard, there might come a point where you have to acknowledge that local LLMs are not for you. None of this is really that hard with some basic Google bootstrap skills or by asking an LLM to help with the command.

                                                            • kccqzy

                                                              today at 5:44 PM

                                                              There are easier ways to run it. OP seemed to enjoy tinkering and customizing the command to run it exactly the way they want. When I don’t want to tinker Unsloth Studio is probably closest to pick a model and voila.

                                                              • losthubble

                                                                today at 5:29 PM

                                                                just tell claude/codex "set this up on my system $huggingfacelink"

                                                                • xienze

                                                                  today at 4:55 PM

                                                                  Well there's a lot of knobs to turn if you want to improve performance. You can always point an LLM at the model card, give it your info, and have it write up the command.

                                                                    • Auracle

                                                                      today at 4:58 PM

                                                                      Sure, but shouldn’t the programs to run the LLMs go “the user has this much vram and the model is this size, so I’ll start with sensible defaults based on that”?

                                                                      You could override, obviously.

                                                                        • zargon

                                                                          today at 5:05 PM

                                                                          Yes, llama.cpp does that.

                                                                  • skrebbel

                                                                    today at 5:05 PM

                                                                    > LM Studio doesn't work behind proxies.

                                                                    Woa, is that still a thing? You mean like SOCKS5 stuff that you have to manually configure in every application that uses the internet?

                                                                    I mean maybe I'm just living under a rock but I feel like that's a rather niche situation you got there.

                                                                      • bmitc

                                                                        today at 5:07 PM

                                                                        > I feel like that's a rather niche situation you got there

                                                                        Every big company in the world uses a network proxy. LM Studio, as far as I can tell, cannot be configured to work behind such proxies.

                                                                          • Aurornis

                                                                            today at 5:41 PM

                                                                            > Every big company in the world uses a network proxy.

                                                                            It's becoming more rare, now.

                                                                            A lot of the universal truths about corporate networks from the early 2000s are no longer true today. Some companies are stuck in their ways though.

                                                                            The overlap between companies that require someone to use a network proxy and companies that have GPU-equipped machines with enough RAM for LLMs and and that allow people to download and run executables of their choosing has to be small.

                                                                            • vardump

                                                                              today at 5:37 PM

                                                                              Every big company? YMMV, but I'd say about 20-40% do.

                                                                              • skrebbel

                                                                                today at 5:10 PM

                                                                                Woa TIL. I thought that was somehow long solved at the OS level or with VPNs or something like that (no idea exactly how, I'm sure just I'm misunderstanding something basic).

                                                                                Makes it rather weird that LM Studio doesn't support it given how their target market, or well at least for their paid products, is very enterprisey.

                                                                                • ThreatSystems

                                                                                  today at 5:38 PM

                                                                                  If you're on Linux you can probably use proxychains.

                                                                          • naasking

                                                                            today at 5:03 PM

                                                                            You know free LLMs can help you understand that command line or design your own...

                                                                            • CamperBob2

                                                                              today at 4:59 PM

                                                                              I never install this stuff manually anymore. Just tell your LLM of choice to download model X from URL Y, build the latest inference engine of choice E, and then create batch files or shell scripts to run instruct and/or reasoning models in accordance with instructions at URL Z.

                                                                      • KronisLV

                                                                        today at 3:12 PM

                                                                        I hope really badly that we'll get a new 35B A3B or similar MoE model!

                                                                        I also miss the Qwen 3 Coder Next, which was 80B A3B, there are quite a few use cases where a non-dense model <100B would be the sweet spot (when you have the VRAM but not the TDP or compute power). Heck, I'd gladly take A5B or A8B or even A10B as a sort of middle ground.

                                                                        Also alternate link for viewing the images without signing in: https://xcancel.com/Alibaba_Qwen/status/2088280182356611304

                                                                          • Casteil

                                                                            today at 3:19 PM

                                                                            I'm hoping too that they'll put out some MoE variants.

                                                                            Qwen3.5:122b:a10b can run about twice as fast as this 27b dense model.

                                                                            Edit: Like its predecessors, 3.8 seems really inclined to overthinking, and on a 27b dense model that's kind of painful. I think I'm going to stick with gemma4:26b-a3b as my go-to because it runs about 4x as fast and tends to only need a fraction of the tokens in its 'thinking' stage to get the same or similar answer.

                                                                              • satvikpendem

                                                                                today at 5:04 PM

                                                                                Reduce or turn down thinking:

                                                                                https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates

                                                                                  • Casteil

                                                                                    today at 5:37 PM

                                                                                    Yeah, that's probably the answer given that it apparently defaults to 'xhigh'.

                                                                                      • dannyw

                                                                                        today at 5:44 PM

                                                                                        Probably helps it score a little bit better in benchmarks :) `medium` seems like a nice balance so far; along with some light steering to vary think effort as needed for task and being pragmatic.

                                                                                • Phemist

                                                                                  today at 4:50 PM

                                                                                  Did you try the claude reasoning traces finetune for qwen3.6? I find that it works muuch better. I assume the same 3.8 finetune will be released at some pointas well.

                                                                                  Edit: link - https://huggingface.co/rico03/Qwen3.6-27B-Claude-Opus-Reason...

                                                                              • peri-cl

                                                                                today at 3:20 PM

                                                                                Same here! Qwen3.6-35B-A3B is the only local model I've found that runs reasonably on my iGPU. Looks like me and and my noisily-wheezing laptop will be sitting out this upgrade.

                                                                                  • expedited123

                                                                                    today at 4:44 PM

                                                                                    Mind sharing your laptops specs? Just interested to see what is needed to locally run Qwen3.6-35B-A3B

                                                                                      • peri-cl

                                                                                        today at 5:11 PM

                                                                                        Don't mind! It's 64 GiB dual-channel DDR5-6400, i.e. roughly 100 GiB/s of bandwidth. (AMD 7840U (Zen 4))

                                                                                        I'm using a Q4 quantization from unsloth (Qwen3.6-35B-A3B-UD-Q4_K_XL). It gets up to ~20 tokens/second in generation. I don't know precisely how much KV cache I can safely use, but it's in between 140k–256k. (I.e., 140k reliably works, 256k kernel-crashes from OOM. Don't feel like bisecting).

                                                                                        [edit to add: If anyone's curious, I've now tested the new, non-MoE, Qwen3.8-27B: ~4 token/s second. (Qwen3.8-27B-UD-Q4_K_XL). This is why I'm sticking with MoE!]

                                                                                        Inference is llama.cpp with the Vulkan GPU backend on Linux. (I.e., -DGGML_VULKAN=1 on the llama.cpp build, and --gpu-layers all on llama-cli or llama-server. (And for my specific setup, two kernel parameters specific to amdgpu: ttm.pages_limit and ttm.page_pool_size. A driver VRAM limiter. Look it up if you're on amdgpu!)).

                                                                                          • expedited123

                                                                                            today at 5:49 PM

                                                                                            Thanks! I don't have any knowledge of running models locally.

                                                                                            I assume it would not be able to handle an unquantized Qwen3.6-35B or is it irrelevant as you almost always would want to run a quantized version of the model on consumer hardware?

                                                                                              • seanmcdirmid

                                                                                                today at 5:51 PM

                                                                                                not parent, but 4-bit quantization is generally consider a good trade off for speed/performance, so you might use it even when you aren't on consumer hardware, but definitely when you are on consumer hardware.

                                                                                    • cyanydeez

                                                                                      today at 5:21 PM

                                                                                      yeah, that's the A3B part; going up to A5B would probably also feel comfortable.

                                                                                      on the 395+ AI MAX w/128GB, the A10B qwen 3.5 can do a lot of long running work if you don't need to baby sit it. deer-flow works well like that.

                                                                                  • jwr

                                                                                    today at 4:42 PM

                                                                                    Me too. 35B A3B runs really fast on my MacBook Pro (M4 Max) and is suitable for real-time tasks like dictation post-processing. The dense model is not.

                                                                                    • Alifatisk

                                                                                      today at 3:40 PM

                                                                                      > I'd gladly take A5B or A8B or even A10B as a sort of middle ground.

                                                                                      Whats up with focusing on the active param count? Do yall fiddle with the weights or something?

                                                                                        • kennywinker

                                                                                          today at 4:28 PM

                                                                                          Total param count decides how much vram you need to run it. Active param count decides how fast it runs. My 10 year old GPU can load quantized 35B or 27B, but it can’t process 27B parameters per token faster than 2-4tok/s, while it can do A3B at >40tok/s

                                                                                            • Alifatisk

                                                                                              today at 4:53 PM

                                                                                              Thank you Kenny

                                                                                          • martinald

                                                                                            today at 3:44 PM

                                                                                            You can run these on CPUs at a somewhat reasonable speed.

                                                                                              • KronisLV

                                                                                                today at 4:05 PM

                                                                                                Or (somewhat) low TDP GPUs for that matter, like workstation ones, that might have enough total VRAM but not the best bandwidth/compute.

                                                                                    • scrlk

                                                                                      today at 3:11 PM

                                                                                      Beats Opus 4.7 Max (w/ Claude Code) on DeepSWE (42.2 vs 40). Looks like Qwen's 27B models continue to pack some punch.

                                                                                      Unsloth's GGUF quants are up: https://huggingface.co/unsloth/Qwen3.8-27B-GGUF

                                                                                        • NitpickLawyer

                                                                                          today at 3:25 PM

                                                                                          > Beats Opus 4.7 Max

                                                                                          I'm a huge open model fan, and have used them since forever, even have daily drivers for on-prem dev, but no. They do not beat opus on real-world usage.

                                                                                          Qwen models are impressively good for what they are, are "good enough" for plenty tasks, can be ran locally on decently priced hardware, and so on. They certainly have their uses, and the field in general has advanced faster than my early expectations. But to compare a 27B model to SotA behemoths from a few months ago is doing everyone a disservice, especially people who pick it up, try to use them just like API models, and leave disappointed and confused. Number goes up on a benchmark isn't it.

                                                                                            • spmurrayzzz

                                                                                              today at 3:34 PM

                                                                                              > They do not beat opus on real-world usage

                                                                                              We have an internal eval that measures performance on tasks for a handful of embedded systems repos for our mmWave radios (mostly Rust, some C for microcontroller stuff). Qwen3.6-27B scores only 4% lower for pass@1, n=250 compared to Opus-4.8.

                                                                                              For the labeled dataset, the average PR size they're being measured against is around 1.5k SLOC.

                                                                                              This is very much "real-world usage" for us. The sort of change sets that come in daily/weekly and are solving non-trivial issues in the respective codebases.

                                                                                              As is usually the case, the most broad claims from both the labs and from the consequent pushback are talking past each other.

                                                                                                • croemer

                                                                                                  today at 4:58 PM

                                                                                                  How much does it score though? 0% would be 4% less if Opus was at 4%. Unless you mean relative fraction not percentage points - but people usually mean percentage points in such situations.

                                                                                                    • tyre

                                                                                                      today at 5:34 PM

                                                                                                      0% is not 4% less than 4%, that would be 3.84%.

                                                                                                      0% is 4 percentage points (pp) less than 4%.

                                                                                                        • naikrovek

                                                                                                          today at 6:07 PM

                                                                                                          [dead]

                                                                                                  • cyanydeez

                                                                                                    today at 4:36 PM

                                                                                                    Let us know when you have Qwen vs Qwen comparison stats. As long as there's not a regression, that'd be awesome.

                                                                                                      • spmurrayzzz

                                                                                                        today at 4:39 PM

                                                                                                        4% is within the margin of error anyways for pass@1, so I think pass@k > 1 is gonna be the better indicator of any movement (still need to calibrate the optimal k to re-test). 10 seems too tolerant even though that tends to be the next tranche I reach for.

                                                                                                          • croemer

                                                                                                            today at 4:59 PM

                                                                                                            Depends on where you sit on the binomial curve. At p=0.04 for n=250 4% points would not be within margin of error.

                                                                                                              • spmurrayzzz

                                                                                                                today at 5:35 PM

                                                                                                                [dead]

                                                                                                    • HonshinM

                                                                                                      today at 6:02 PM

                                                                                                      [dead]

                                                                                                      • enraged_camel

                                                                                                        today at 5:26 PM

                                                                                                        >> We have an internal eval that measures performance on tasks for a handful of embedded systems repos for our mmWave radios

                                                                                                        Okay but the parent said real-world usage, presumably meaning coding tasks.

                                                                                                        We have a whole bunch of complex evals that Haiku 4.5 passes. That doesn't mean it is a good model for coding.

                                                                                                          • spmurrayzzz

                                                                                                            today at 5:36 PM

                                                                                                            Yes, these are coding tasks in the embedded systems domain (I mentioned Rust and C).

                                                                                                    • KronisLV

                                                                                                      today at 3:32 PM

                                                                                                      > ...but no. They do not beat opus on real-world usage.

                                                                                                      I agree, but then we just need meaningful benchmarks that clearly show that! Otherwise it's hand waving about something that should be put on paper in quantifiable terms.

                                                                                                        • pimeys

                                                                                                          today at 4:01 PM

                                                                                                          If you are working in a company and using language models, it is a very good idea to hold a bunch of evals you can trust and use to validate new models. Calibrate every once in a while with prod data. We have our own and the only numbers on quality and cost I trust come from this setup.

                                                                                                          • niek_pas

                                                                                                            today at 3:37 PM

                                                                                                            A wise man once said, "not everything that counts can be counted, and not everything that can be counted counts".

                                                                                                            • mlmonkey

                                                                                                              today at 4:19 PM

                                                                                                              In the end, the only benchmark that matters is your own.

                                                                                                              • bewareofscams

                                                                                                                today at 3:38 PM

                                                                                                                Only useful benchmarks are those you (in particular) don't have access to.

                                                                                                                  • rhdunn

                                                                                                                    today at 4:52 PM

                                                                                                                    The only useful benchmarks are those you've created for your specific workflow. Only then can you assess whether a given model is better or worse for what you are using it for.

                                                                                                                    There are tools like promptfoo designed for this.

                                                                                                                • xienze

                                                                                                                  today at 3:41 PM

                                                                                                                  > but then we just need meaningful benchmarks that clearly show that!

                                                                                                                  That's the rub. AI benchmarks are IMO, by and large totally unreliable. We think of them as similar to traditional benchmarks of deterministic processes where the number of variables is low. But they're anything but that. Non-deterministic processes with an astounding number of variables and fuzzy acceptance criteria.

                                                                                                                  It leads to results like these, where if you take it at face value, the only conclusion you can draw is "wow Anthropic must be stupid if Opus takes 1T parameters to do what Qwen can do in 27B."

                                                                                                              • metadat

                                                                                                                today at 4:34 PM

                                                                                                                How can you say this when you haven't even tried it yet? Is it just hypothetical vibes?

                                                                                                                • redox99

                                                                                                                  today at 4:57 PM

                                                                                                                  Yep. These small models are actually worse than GPT 3.5 at some tasks (like recalling facts). You can definitely make models smarter at specific tasks (like tool calling, coding) but you can't compress the entire human knowledge into a 30GB file. It's just not enough bits.

                                                                                                                    • ferrouswheel

                                                                                                                      today at 5:51 PM

                                                                                                                      But why would you use a model to store factual knowledge, that is stupid. We want intelligence, not a database.

                                                                                                                      • ycui7

                                                                                                                        today at 5:25 PM

                                                                                                                        that is why we enable web search for the agent. the memory can come from the internet.

                                                                                                                        deepseek-v4-flash needs web search to return true facts.

                                                                                                                  • willcmcc

                                                                                                                    today at 4:28 PM

                                                                                                                    There is 0 shot you can make that claim about this model you have not used or downloaded yet

                                                                                                                    • altmanaltman

                                                                                                                      today at 4:38 PM

                                                                                                                      "Benchmark is stupid" and "model beats model on benchmark" are two different things, though. The second one is objectively true regardless of your views on the first one, right? To expect everyone to share your opinion that benchmarks are stupid is pretty weird, and just saying "no" to an objective truth is the definition of delusion.

                                                                                                                        • kennywinker

                                                                                                                          today at 5:56 PM

                                                                                                                          If a benchmark is a measure of nothing useful, then model beats model is an objectively useless fact

                                                                                                                  • jrflo

                                                                                                                    today at 4:59 PM

                                                                                                                    That kind of result makes me suspicious of benchmaxxing. Qwen 27B is 100x smaller than Opus 4.7. Is it really 100x more parameter-efficient? Two orders of magnitude is hard to believe. I don't have the hardware to run a 27B, but I'm curious what real world use is like. Maybe I'll have to buy some usage on a cloud provider to run my own tests, but this seems fishy to me.

                                                                                                                      • CuriouslyC

                                                                                                                        today at 5:17 PM

                                                                                                                        Qwen small models are heavily coding focused, whereas Opus is everything to everybody (even if code is their bread and butter). The downside is they'll frequently hallucinate world knowledge so they need to be RL'd to double check their knowledge against sources and verify facts/library names/etc.

                                                                                                                        • dannyw

                                                                                                                          today at 5:47 PM

                                                                                                                          It's very agentic coding focused; and I'd say a good executor but certainly not Opus in scale; overall knowledge; long-horizon work and recovery; etc.

                                                                                                                          e.g. If you try to chat to it about something philosophical for example, or maybe a debate / creative writing, then you'll very quickly see how it is still a much smaller model at the end of the day.

                                                                                                                          Still, it's such a relatively accessible model to run, and I find a big part of leveraging smaller models is to give it well-scoped tasks; not too high level or ambitious ones. Very impressive for its size and the ability to run locally :)

                                                                                                                      • nblgbg

                                                                                                                        today at 3:14 PM

                                                                                                                        Is there any advantage to using the model from Unsloth compared with https://huggingface.co/Qwen/Qwen3.8-27B-FP8 ?

                                                                                                                          • benxh

                                                                                                                            today at 3:15 PM

                                                                                                                            Depends on what software/hardware you'll run it. GGUFs from Unsloth can run on pretty much every single potato; full weights need beefy gpus

                                                                                                                            • petu

                                                                                                                              today at 3:28 PM

                                                                                                                              Unsloth one is gguf for llama.cpp (and some other on-device engines).

                                                                                                                              So advantage is not having to produce your own quantisation / gguf from .safetensors you've linked.

                                                                                                                              • danielhanchen

                                                                                                                                today at 4:04 PM

                                                                                                                                We also made NVFP4 ones if that helps! https://huggingface.co/unsloth/Qwen3.8-27B-NVFP4

                                                                                                                                  • hadlock

                                                                                                                                    today at 5:55 PM

                                                                                                                                    This is the version we'll be testing on our rtx 6000 today! Thank you

                                                                                                                                • satvikpendem

                                                                                                                                  today at 5:10 PM

                                                                                                                                  Unsloth usually also fixes the models when they bork something, which always happens. For Gemma for example the tool calling wasn't working for the longest time.

                                                                                                                                    • danielhanchen

                                                                                                                                      today at 5:22 PM

                                                                                                                                      That wasn't our problem right? Gemma officially updated tool calling which we adopted

                                                                                                                                  • ycui7

                                                                                                                                    today at 5:26 PM

                                                                                                                                    if you have the VRAM, use offical release. quantized model lose focus after long context and can do damages or thinking loop

                                                                                                                                    • 4chandaily

                                                                                                                                      today at 3:29 PM

                                                                                                                                      Run the unsloth if you are using llama.cpp (GGUF)

                                                                                                                                      Run the one you linked if you are running vllm (safetensors)

                                                                                                                                  • Aurornis

                                                                                                                                    today at 5:32 PM

                                                                                                                                    In the local LLM communities there is a lot of respect for the Qwen models, but everyone comes to acknowledge that they do a lot of benchmaxxing after using them. Even at full precision they're never as good as models with similar benchmarks.

                                                                                                                                    • Foobar8568

                                                                                                                                      today at 3:45 PM

                                                                                                                                      Considering the clusterfuck that is opus 5 or even fable, if Qwen 27B is trully better than Opus 4.7 Max, I will rejoice.

                                                                                                                                        • ferrouswheel

                                                                                                                                          today at 5:53 PM

                                                                                                                                          Yeah Opus 5 is almost unusable as a daily driver without making me go insane from excessive claude babble.

                                                                                                                                          • UncleOxidant

                                                                                                                                            today at 3:59 PM

                                                                                                                                            If it's as good as Sonnet 4.6 for most things I'd be happy.

                                                                                                                                        • WithinReason

                                                                                                                                          today at 3:23 PM

                                                                                                                                          I wish each quant was benchmarked on the same tests as the original network so we could compare their performance

                                                                                                                                            • scrlk

                                                                                                                                              today at 3:48 PM

                                                                                                                                              Unsloth publishes KL divergence numbers which measures how much the quantised probability distribution changes vs unquantised: https://unsloth.ai/docs/models/qwen3.8#quantization-analysis

                                                                                                                                              It's a bit bare at the moment, I assume they are going to add further detail later (eg comparison to other quants), similar to their other releases.

                                                                                                                                                • zargon

                                                                                                                                                  today at 4:36 PM

                                                                                                                                                  KL divergence is nothing close to a replacement for benchmarks. As flawed as benchmarks are, KL divergence is a barely useful signal. The fact that Unsloth only just started publishing KL divergences shows how unserious the quantization space is.

                                                                                                                                                    • xscott

                                                                                                                                                      today at 6:05 PM

                                                                                                                                                      You're very right about KL divergence. I spent a couple days playing with the Gemma 4 models. That's 10 separate models (varying weights, MoE, QAT or not, etc...) with identical tokenizers. I treated 31B at BF16 as the gold standard, feeding Wikipedia snippets, and anthropomorphizing a bit:

                                                                                                                                                      Gemma 4 31B: "Um, if I really said all of that, I guess I'd say this next"

                                                                                                                                                      Gemma 4 26B: "Dude, I would've said completely different stuff" (large divergence)

                                                                                                                                                      Gemma 4 12B: "Umm, there's zero chance I would've said some of this" (INFINITE divergence)

                                                                                                                                                      Gemma 4 E4B and E2B: "Derp derp, I'm happy to say almost anything" (lowest divergence)

                                                                                                                                                      For models which are chat trained, they simply would not recite Wikipedia, so the divergence is almost meaningless. I thought about capturing a realistic coding session and trying to use that as the corpus, but you need to preserve the turn-based tokens and such, so I moved on to other things.

                                                                                                                                                      • danielhanchen

                                                                                                                                                        today at 5:24 PM

                                                                                                                                                        Actually we do publish non KLD benchmarks - top-1% is better - for NVFP4 for eg we did MMLU Pro, GPQA, AIME 2025: https://unsloth.ai/docs/models/qwen3.6#nvfp4-benchmarks

                                                                                                                                                        Sometimes they're just slow and expensive, so we we KLD as a proxy measure and it's very high correlation (95%+)

                                                                                                                                                        • lostmsu

                                                                                                                                                          today at 5:02 PM

                                                                                                                                                          > The fact that Unsloth only just started publishing KL divergences shows how unserious the quantization space is.

                                                                                                                                                          Just wanted to say that this is a very important point that I totally agree with. People are obsessed with KL divergence, but it is yet to be demonstrated to be a descent proxy for agentic coding benchmarks.

                                                                                                                                                      • WithinReason

                                                                                                                                                        today at 5:01 PM

                                                                                                                                                        That's not a replacement for benchmarks

                                                                                                                                                    • cpburns2009

                                                                                                                                                      today at 5:32 PM

                                                                                                                                                      When I tested various eval benchmarks on Qwen3.5/3.6 27B with Unsloth's quants, the scores usually dropped 0-5% between UD-Q6 and UD-Q3 depending on the eval.

                                                                                                                                                  • edg5000

                                                                                                                                                    today at 3:13 PM

                                                                                                                                                    That's crazy, considering the massive size difference. But the small Qwen models are known for punching above their weight.

                                                                                                                                                    • UncleOxidant

                                                                                                                                                      today at 3:57 PM

                                                                                                                                                      Good morning Dario!

                                                                                                                                                  • Balinares

                                                                                                                                                    today at 5:53 PM

                                                                                                                                                    I wonder if Anthropic and OpenAI possibly missed the window to go public. A 27B open-weight model trading blows with the SOTA from just half a year ago is not great news for trillion-dollar investments...

                                                                                                                                                      • brcmthrowaway

                                                                                                                                                        today at 5:54 PM

                                                                                                                                                        This is why they've been making bank on the secondary market. They can retire now.

                                                                                                                                                    • ramon156

                                                                                                                                                      today at 3:42 PM

                                                                                                                                                      People will claim it's not comparable to Opus despite it beating the score. I'm not sure I disagree, but I'm also unsure whether I care. Most new models nowadays are "good enough". I cannot complain because I'd rather spend that time improving my prompts and docs. Opus might be a _slight bit better_ at picking up vague hints, but it's also extremely expensive, and I hit the 5 hour limit way too quick.

                                                                                                                                                      I care a lot about speed and efficiency right now. For my setup I would like to have 2-3 different model families. I've settled on GLM-5.3 (formerly Deepseek v4 pro 0813) for architecting, Deepseek V4 Pro 0813 for developing, and Gemini flash lite (any recent cheap model) for repo scouting. I'll add another one in the mix for reviewing (in this case Gemini 3.7) and that's all I need.

                                                                                                                                                      I've tried most models except Grok.

                                                                                                                                                      Qwen is too expensive IMO (Alibaba Cloud subscriptions are hard to come by and I'm not spending 50 euros a month for a tool, so 18 euros it is). If it ever becomes efficient enough to run locally I will definitely look back.

                                                                                                                                                      Claude is slow and expensive (the cache hit prices are absurd).

                                                                                                                                                      OAI is pretty good, I might add it to my arsenal seeing how cheap it is.

                                                                                                                                                      These opinions change every day. Last week I would've never picked Deepseek until I read about the pricing. even post aug 16 it's worth it (although it's getting close to gemini pricing).

                                                                                                                                                      Right now my costs are 12 euros a month (z.ai) + whatever deepseek consumes. This typically isn't more than 8 euros a week. 44 euros a month and I have a setup that is doing pretty well.

                                                                                                                                                        • satvikpendem

                                                                                                                                                          today at 5:12 PM

                                                                                                                                                          You should check out Grok, it's quite a good deal from the Cursor subscription side but it's cheap even by API prices.

                                                                                                                                                            • altruios

                                                                                                                                                              today at 5:44 PM

                                                                                                                                                              No serious person or sane person uses the LLM that's constantly being tweaked by an anti-woke white-genocide-supporting weird little man. Don't feed the totalitarian wannabe's (or the totalitarians in general, for that matter).

                                                                                                                                                                • satvikpendem

                                                                                                                                                                  today at 5:50 PM

                                                                                                                                                                  Yeah I'm sure everyone on r/cursor or in previous HN threads about Grok 4.5 or 4.6 are all unserious and insane.

                                                                                                                                                                  No one actually cares about the politics as long as the model codes well.

                                                                                                                                                                    • kennywinker

                                                                                                                                                                      today at 6:03 PM

                                                                                                                                                                      This is the “Mussolini made the trains run on time” of ai hot takes.

                                                                                                                                                                      (Btw, mussolini didn’t make the trains run on time)

                                                                                                                                                              • ai_fry_ur_brain

                                                                                                                                                                today at 5:18 PM

                                                                                                                                                                [dead]

                                                                                                                                                            • hypfer

                                                                                                                                                              today at 3:45 PM

                                                                                                                                                              > I've settled on GLM-5.3 (formerly Deepseek v4 pro 0813) for architecting

                                                                                                                                                              Dude, GLM-5.3 released _today_.

                                                                                                                                                              The phrasing "I've settled on" is incorrect for this context.

                                                                                                                                                                • ramon156

                                                                                                                                                                  today at 3:46 PM

                                                                                                                                                                  hence the "former deepseek v4 pro". I tried it out this morning and have had no complaints. I already liked glm 5.2

                                                                                                                                                                    • kristjansson

                                                                                                                                                                      today at 5:29 PM

                                                                                                                                                                      > Deepseek v4 pro 0813

                                                                                                                                                                      Which itself released yesterday? You're writing, reading, and evaluating enough software in a ~36 hour period to form, reject, and form another opinion about which model makes better _architectural_ choices?

                                                                                                                                                                      • Topfi

                                                                                                                                                                        today at 4:41 PM

                                                                                                                                                                        Honest question, how do you assess models this quickly? What metrics are you using? Would love to get my suite from multiple days and hundreds of prompts down to minutes. Got a few first pass tasks I run upon release for an initial experience, but those only work because even Fable and Sol fail despite objectively correct solutions existing, so it works because most models fail, but then, those are consciously not enough for coding, tool use, adherence or task specific inference and assessment…

                                                                                                                                                                          • satvikpendem

                                                                                                                                                                            today at 5:11 PM

                                                                                                                                                                            What are you working on? That can dictate which models are best.

                                                                                                                                                                        • hypfer

                                                                                                                                                                          today at 3:48 PM

                                                                                                                                                                          The sentence still doesn't make sense, because "settled on" implies a long testing phase with a verdict eventually emerging out of that.

                                                                                                                                                                          What you're currently doing is "testing out"

                                                                                                                                                                  • tosh

                                                                                                                                                                    today at 4:10 PM

                                                                                                                                                                    i think you will like luna if you haven't tried it yet

                                                                                                                                                                    • simplyluke

                                                                                                                                                                      today at 4:06 PM

                                                                                                                                                                      I'm convinced a lot of the anti-open-weight model comments at this point are inorganic traffic - there's trillions in investor money riding on a world where these models aren't cheap commodities. Having actually used things like the recent GLM, Kimi, and Qwen I think any edge the labs have is marginal at most and actually prefer the open weight models in most day to day usage.

                                                                                                                                                                      Anthropic's recent releases are wordy to the point of exhaustion. Every time I use opus recently I find myself wanting to yell "GET TO THE POINT" at a terminal, which is exacerbated by it being slow.

                                                                                                                                                                  • satvikpendem

                                                                                                                                                                    today at 5:05 PM

                                                                                                                                                                    As usual, the Jinja templates are messed up so use this [0] to reduce or turn off thinking, fix tool calling, keep a 100% KV cache hit rate, etc.

                                                                                                                                                                    [0] https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates

                                                                                                                                                                      • skrebbel

                                                                                                                                                                        today at 5:14 PM

                                                                                                                                                                        I'd love to understand this more. Are you saying the Qwen team spends their very impressive human and compute resources on publishing these amazing models and then botches the chat template with mundane bugs?

                                                                                                                                                                        Like maybe I just misunderstand what's the hard part but wouldn't you assume that people who can put together an impressive model can also write a proper jinja chat template for it?

                                                                                                                                                                          • satvikpendem

                                                                                                                                                                            today at 5:52 PM

                                                                                                                                                                            Yes for the first question. Google of all companies didn't even get it right with Gemma for a while until recently. For some reason it doesn't seem like people can actually get these templates right.

                                                                                                                                                                            • Der_Einzige

                                                                                                                                                                              today at 5:16 PM

                                                                                                                                                                              Yes yes, oh god yes. They also spread FUD in the form of terrible recommended sampler settings.

                                                                                                                                                                              If you're using llamacpp, turn on top-n-sigma with sigma of 1, turn off top-p/top-k. You'll thank me later.

                                                                                                                                                                                • skrebbel

                                                                                                                                                                                  today at 5:34 PM

                                                                                                                                                                                  How is terrible settings a case of FUD?

                                                                                                                                                                      • onlyrealcuzzo

                                                                                                                                                                        today at 3:16 PM

                                                                                                                                                                        If the benchmarks don't lie, this is getting very close to Opus 4.6 capability - which was the turning point for me for when AI was "good enough" that it became very hard to justify not using it.

                                                                                                                                                                        I'm sure there's some benchmaxxing going on, and some things you get only with a a larger model.

                                                                                                                                                                        But I'm feeling pretty confident if not by Gemma 5 than by mid 2028 we'll have local models that are almost always as good as Opus 4.6 was and in many cases far better.

                                                                                                                                                                          • DanielHB

                                                                                                                                                                            today at 3:36 PM

                                                                                                                                                                            What kind of things you only get with a larger model?

                                                                                                                                                                              • onlyrealcuzzo

                                                                                                                                                                                today at 4:47 PM

                                                                                                                                                                                Similar to the way they asked Sol to solve Erdos problems, that's what I want my model to do for programming.

                                                                                                                                                                                I don't want to try to take my best educated guess at what the best design is BEFORE implementation - especially if you're designing a feature for a codebase you're not an expert in, you don't know like the back of your hand (i.e. one that is mostly or entirely LLM generated).

                                                                                                                                                                                What sounds good on paper - often times becomes unideal in practice when you get to the reality of implementation.

                                                                                                                                                                                It may not be worth re-architecting your entire system to get to a "pure" design that would be the best - all things considered.

                                                                                                                                                                                Instead, I'd like the model to independently design many plausible and coherent good solutions, then implement each of them, then intelligently pick the few winners (after its fixed any bugs that could be causing promising solutions to look artificially bad) - unless there's an obvious one - and then give me the data I need to make an informed decision on which one to go with, all before I even look at the design or implementation.

                                                                                                                                                                                You're not getting this from a one shot prompt from a 30B model today. You can't even really get it from Sol or Fable - IME. But you can get somewhat close.

                                                                                                                                                                                  • ferrouswheel

                                                                                                                                                                                    today at 6:01 PM

                                                                                                                                                                                    You're going to be waiting for a while.

                                                                                                                                                                                    Even Fable is bad at this, I would constantly have to fix it going down architectural dead ends or just making obvious mistakes.

                                                                                                                                                                                    Which sucks for people that want LLMs to do everything like a genie, but does mean senior engineers have a few more years before they become redundant.

                                                                                                                                                                                    • alex7o

                                                                                                                                                                                      today at 4:50 PM

                                                                                                                                                                                      This is a harness problem not a model problem, try prime agent it can do that and it will do it well even :P but you need to prompt it in according to its tools and processes.

                                                                                                                                                                                  • redox99

                                                                                                                                                                                    today at 5:06 PM

                                                                                                                                                                                    Asking it factual information[1]. You just can't compress the entire human knowledge into a 30GB file.

                                                                                                                                                                                    [1] Without searching the internet. And even if you allow it, you'll get much worse results because search means browsing and parsing the top results, and search results are horrible, whereas internal knowledge from training encompasses the entire internet plus all books including very niche stuff.

                                                                                                                                                                                    • versteegen

                                                                                                                                                                                      today at 3:46 PM

                                                                                                                                                                                      IME using 5.6 Luna and DS V4 Flash, I notice that although they are excellent at programming, even Opus-like in the way they try to debug, the thing they are worst at is inferring user intent and making good decisions with little information. They are absolutely terrible at that, will misinterpret small wording ambiguities. I suspect that's an ability you can't add with RL training, that it requires the depth of understanding from vast pre-training.

                                                                                                                                                                              • LeBit

                                                                                                                                                                                today at 3:24 PM

                                                                                                                                                                                https://xcancel.com/Alibaba_Qwen/status/2088280182356611304

                                                                                                                                                                                  • looksjjhg

                                                                                                                                                                                    today at 3:34 PM

                                                                                                                                                                                    I could kiss you right now

                                                                                                                                                                                      • newaccount670

                                                                                                                                                                                        today at 3:42 PM

                                                                                                                                                                                        [dead]

                                                                                                                                                                                • kanemcgrath

                                                                                                                                                                                  today at 5:55 PM

                                                                                                                                                                                  I think I am going to buy a second rtx 3060, as 27B has been just outside of my range for to long, and this looks like the parameter count tipping point

                                                                                                                                                                                  • Casteil

                                                                                                                                                                                    today at 4:04 PM

                                                                                                                                                                                    One thing a lot of people don't seem to factor when hyping Qwen is how much models like this tend to 'overthink' with seemingly endless 'second guessing'. 3.8 seems no different from what I've tried thus far.

                                                                                                                                                                                    As capable as it is, it's hard to justify using it when a competing model (e.g. Gemma4:26b-a3b) can consistently achieve the same or similar response with only 1/10th as many 'thinking' tokens, achieve much higher tokens/second, and take a small fraction of the time. I suppose 'YMMV' depending on your use case.

                                                                                                                                                                                    Also, I haven't used it enough yet to see if it's prone to infinite looping, but its predecessors sure were.

                                                                                                                                                                                      • satvikpendem

                                                                                                                                                                                        today at 5:14 PM

                                                                                                                                                                                        Reduce or turn off thinking:

                                                                                                                                                                                        https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates

                                                                                                                                                                                          • Casteil

                                                                                                                                                                                            today at 5:44 PM

                                                                                                                                                                                            Given that it apparently defaults to 'xhigh', this is probably the answer.

                                                                                                                                                                                            Granted, it's still much lower tokens/s than you'll get out of many MoE models.

                                                                                                                                                                                            • IronWolve

                                                                                                                                                                                              today at 5:41 PM

                                                                                                                                                                                              Thank you, this is exactly what I needed.

                                                                                                                                                                                          • lrvick

                                                                                                                                                                                            today at 4:08 PM

                                                                                                                                                                                            Use 3.6 27b as a daily driver for months with charmbracelet crush. Gemma 26b-A3b is not even remotely comparable in terms of coding for me. YMMV depending on how you work, what harness you use, etc I suppose.

                                                                                                                                                                                            • ThouYS

                                                                                                                                                                                              today at 4:14 PM

                                                                                                                                                                                              gemma4 can't hold a candle to 3.6

                                                                                                                                                                                              • cyanydeez

                                                                                                                                                                                                today at 4:16 PM

                                                                                                                                                                                                You can add a thinking budget thats not much effort in llamacpp. You can align the cut off message with your agent instructions.

                                                                                                                                                                                                What you describe is a engineering harness problem.

                                                                                                                                                                                                If you, and i mean the royal you, actually read tge thinking traces you can see and figure out where its stuck

                                                                                                                                                                                                This means an effective harness would observe when the model is overthinking and step in with reasonable redirection, like increasing logging.

                                                                                                                                                                                                Llamacpp can set reasoning budget and message per reauest, so it can be dynamic.

                                                                                                                                                                                                Your complaint is "skill issue" based and will be resolved by people who do something ither than vibe code react demos.

                                                                                                                                                                                            • jedbrooke

                                                                                                                                                                                              today at 3:21 PM

                                                                                                                                                                                              I hope the bonsai team makes another 1bit quant of this model (or releases code/instructions on how to do it), using the Qwen3.6 27B on my 16GB mac mini has been wild . The 1bit quant feels like opus level… for the first couple turns. Then it has trouble eg switching from plan mode to act mode. This is mostly mitigated by starting a new session. (tbf this limitation is called out on the hf page)

                                                                                                                                                                                              I saw unsloth has 1bit quants too so I might check that out, anybody have experience with those?

                                                                                                                                                                                                • spwa4

                                                                                                                                                                                                  today at 4:24 PM

                                                                                                                                                                                                  Sounds like you need to check what the max context is set to ...

                                                                                                                                                                                                    • jedbrooke

                                                                                                                                                                                                      today at 4:40 PM

                                                                                                                                                                                                      100k is all the context I have ram for, this is with any auto-compact turned off. This is using Cline in vs code. I’m sure I could tune the system prompt and mode switching more to work better with this specific model, but I haven’t gone down the custom harness rabbit hole yet.

                                                                                                                                                                                                      And this is also specifically for the 1bit quant version. I don’t think the fp8 or even fp4 versions have this issue, but I haven’t tried those much

                                                                                                                                                                                              • esotericsean

                                                                                                                                                                                                today at 5:57 PM

                                                                                                                                                                                                Need to upgrade to a second 3090! Slowly building up my local models with Krea2, MiniMax H3 (and their new Music3), and now Qwen 3.8

                                                                                                                                                                                                • T0mSIlver

                                                                                                                                                                                                  today at 3:48 PM

                                                                                                                                                                                                  Unsloth Q4_K_M on a single 3090, llama.cpp "Generate an SVG of a pelican riding a bicycle" first try https://www.reddit.com/r/LocalLLaMA/comments/1voa3ch/comment...

                                                                                                                                                                                                    • btbuildem

                                                                                                                                                                                                      today at 6:01 PM

                                                                                                                                                                                                      Most people cannot draw a bicycle that well!

                                                                                                                                                                                                  • literoldolphin

                                                                                                                                                                                                    today at 5:33 PM

                                                                                                                                                                                                    Why is anyone even using video cards these days? You may as well be burning cash.

                                                                                                                                                                                                    This is the perfect candidate for just splattering it on your nvme and then reading it off there and into memory. All of these run perfectly fine on simple m4 silicone:

                                                                                                                                                                                                    https://github.com/drumih/turbo-fieldfare

                                                                                                                                                                                                    https://github.com/leonickson1/Swiftlet

                                                                                                                                                                                                    https://github.com/sqliteai/warp

                                                                                                                                                                                                      • awkwardpotato

                                                                                                                                                                                                        today at 5:38 PM

                                                                                                                                                                                                        Those are all for MoE models. And I prefer measuring my tokens in t/s instead of s/t

                                                                                                                                                                                                        • ferrouswheel

                                                                                                                                                                                                          today at 6:03 PM

                                                                                                                                                                                                          Lol, "burn money on apple hardware instead!"

                                                                                                                                                                                                      • Almondsetat

                                                                                                                                                                                                        today at 4:52 PM

                                                                                                                                                                                                        The $1500 Intel B70 with 32GB of VRAM can run this model at max context with good performance, btw. If you don't want to drop $5-10k for running DeepSeek this is your best budget option for local refactor/small scale dev help

                                                                                                                                                                                                          • LeBit

                                                                                                                                                                                                            today at 5:19 PM

                                                                                                                                                                                                            I understand the B70 is a bargain vs AMD and especially nVidia offerings, but to me it feels like I would be buying something that would feel too limited in less than a year. 48G would be much more confortable.

                                                                                                                                                                                                            And I know the 96G nVidia cards are selling for over 10k$.

                                                                                                                                                                                                            The future can’t arrive fast enough!

                                                                                                                                                                                                              • Almondsetat

                                                                                                                                                                                                                today at 6:06 PM

                                                                                                                                                                                                                32GB is perfect for models around 30B parameters. Since qwen has really hit the spot with their 27B dense models, I think it's a good bet. Also, 32GB is enough for other tasks such as image/video generation and loading multiple smaller specialized models

                                                                                                                                                                                                            • aappleby

                                                                                                                                                                                                              today at 5:40 PM

                                                                                                                                                                                                              I have a B70, what llama options are you using and what performance are you seeing?

                                                                                                                                                                                                              • segmondy

                                                                                                                                                                                                                today at 5:23 PM

                                                                                                                                                                                                                You don't need $10k to run DeepSeek, I run it on a $1000 system.

                                                                                                                                                                                                                  • 758488

                                                                                                                                                                                                                    today at 6:03 PM

                                                                                                                                                                                                                    Could you elaborate please? Genuinely interested

                                                                                                                                                                                                                • bogzz

                                                                                                                                                                                                                  today at 4:53 PM

                                                                                                                                                                                                                  Oh, can it work with the /v1/completions/ auto-complete endpoint?

                                                                                                                                                                                                                    • Almondsetat

                                                                                                                                                                                                                      today at 4:59 PM

                                                                                                                                                                                                                      Sorry, I wrote autocompletion by force of habit. I simply meant it can complete code you have already created a structure for, which personally is very nice

                                                                                                                                                                                                                        • bogzz

                                                                                                                                                                                                                          today at 5:01 PM

                                                                                                                                                                                                                          I thought so, but thanks for the clarification. I am a little bit disappointed that local autocompletion models have been left by the wayside in favor of models post-trained for agentic coding. Both Codestral and Qwen-2.5-coder are more than a year old at this point, but local auto-complete seems to me to be such a great usecase.

                                                                                                                                                                                                              • TomGarden

                                                                                                                                                                                                                today at 3:16 PM

                                                                                                                                                                                                                Any tips on the best approach at running this at an M4 Max 128GB? Token throughput was a bit slow with the last 27B one (MLX), ended up using the A3B variant but if I could get this one to reasonable speed I'd much prefer it.

                                                                                                                                                                                                                  • mft_

                                                                                                                                                                                                                    today at 3:28 PM

                                                                                                                                                                                                                    Go for a slightly more quantised version, and experiment with different MTP settings. I find that MLX versions are marginally faster on my 64GB M1 Max, but I usually use Unsloth's GGUFs via llama.cpp as there's a much greater range of quants available and I prefer llama.cpp. MTP sometimes also helps a little, but I suspect it's less helpful on my system than others.

                                                                                                                                                                                                                    Unsloth: https://huggingface.co/unsloth/Qwen3.8-27B-GGUF

                                                                                                                                                                                                                    This might work for you, but I didn't get on very well with MTPLX when I tried it a while back; YMMV: https://huggingface.co/Youssofal/Qwen3.8-27B-MTPLX-Optimized...

                                                                                                                                                                                                                      • evgen

                                                                                                                                                                                                                        today at 3:55 PM

                                                                                                                                                                                                                        This is the way if you need speed. It costs a little bit in smarts, but compare the MTPLX option listed above with the oQ4e-mtp quant using oMLX. The good cacheing layer in oMLX will help things feel faster for some classes of tasks in my experience.

                                                                                                                                                                                                                    • jwr

                                                                                                                                                                                                                      today at 4:47 PM

                                                                                                                                                                                                                      I have an M4 Max (unfortunately 64GB). I have been running the Qwen 35B A3B one for a while now, after testing and benchmarking a number of models. That one was consistently the best in class for tasks like despamming, E-mail classification, OCR and dictation post-processing. It was also really fast (90 tokens/s).

                                                                                                                                                                                                                      I'm benchmarking the 3.8 model now, it seems it is better (near-perfect score on my E-mail spam filtering benchmark, best of any model I tested, ever). But it is slow.

                                                                                                                                                                                                                      One thing I would recommend is keeping an eye on MTP parameters. I tested and benchmarked extensively, and I use `--spec-draft-n-max 2` with llama.cpp. Longer sequences actually decrease overall performance.

                                                                                                                                                                                                                      As for running, I ended up using llama.cpp and its llama-server, with a bunch of scripts written by AI, because I got tired of LM Studio not implementing the image-related parameters which made gemma4 useless for OCR.

                                                                                                                                                                                                                      • UncleOxidant

                                                                                                                                                                                                                        today at 4:06 PM

                                                                                                                                                                                                                        Wait for the MTP variants that will likely be out within days. I'm on a 128GB Strix Halo box and for 3.6-27B 8bits I was getting about 9tok/sec (not great). With MTP that gets closer to 18 tok/sec (kind'a usable).

                                                                                                                                                                                                                          • anana_

                                                                                                                                                                                                                            today at 4:20 PM

                                                                                                                                                                                                                            Seems like MTP is available immediately!

                                                                                                                                                                                                                        • seanmcdirmid

                                                                                                                                                                                                                          today at 4:41 PM

                                                                                                                                                                                                                          27B is a dense model so it will be slower with an MoE (A3B), but should have better quality? I still haven’t found very good uses cases on my M3 Max for dense models. Even if you can find a MTP version, it doesn’t help much, especially if you compare against an MoE with MTP as well.

                                                                                                                                                                                                                          • LoganDark

                                                                                                                                                                                                                            today at 3:23 PM

                                                                                                                                                                                                                            Unfortunately, that chip just doesn't really have the memory bandwidth to run this (or nearly any) model at acceptable speeds. I have the exact same chip (M4 Max 128GB) and I've been trying to optimize a completely purpose-built implementation with Fable and this is just not possible. Even if you could reach the full 576GB/s, it's just physically impossible to exceed these numbers with the model's architecture:

                                                                                                                                                                                                                            2 bpw - ~85.7t/s

                                                                                                                                                                                                                            3 bpw - ~58.0t/s

                                                                                                                                                                                                                            4 bpw - ~43.9t/s

                                                                                                                                                                                                                            6 bpw - ~29.5t/s

                                                                                                                                                                                                                            8 bpw - ~22.2t/s

                                                                                                                                                                                                                            16 bpw - ~11.2t/s

                                                                                                                                                                                                                            without cheating. You'd have to exclude layers, skip operations, etc. basically do stuff the model wasn't trained for. And speed collapses so fast with context that even 2 bpw would be looking at ~37.6t/s after just 128K tokens.

                                                                                                                                                                                                                            MTP only improves the situation by up to 2x in the ideal case, while drastically reducing the performance floor. While optimizing a 9B model on this hardware, I've found that the GPU just doesn't have enough FLOPS to handle speculating more than one or two tokens ahead on a single stream, regardless of quant level, simply because of the arithmetic cost of the forward pass. The 27B model would be even more expensive than that, potentially such that it's already bottlenecked by the GPU itself rather than memory.

                                                                                                                                                                                                                            I wouldn't get my hopes up for the 35B-A3B either. Not only is it reportedly much less intelligent, but I hit a similar ~85t/s wall in practice (again with highly specialized inference).

                                                                                                                                                                                                                            Without speculation I can reach around 120t/s on Qwen3.5-9B and with n-gram speculation (not even MTP; this derivative didn't come with one) around about 150t/s on average. This is on the very very edge of what I'd consider acceptable for me to even consider using such a compact model. YMMV due to the silicon lottery but the situation isn't good.

                                                                                                                                                                                                                              • minimaltom

                                                                                                                                                                                                                                today at 5:35 PM

                                                                                                                                                                                                                                What is bpw?

                                                                                                                                                                                                                                Also whats your cutoff for 'acceptable' speed? I would have said 25tok/s.

                                                                                                                                                                                                                                  • rolls-reus

                                                                                                                                                                                                                                    today at 6:07 PM

                                                                                                                                                                                                                                    bits per weight

                                                                                                                                                                                                                            • brcmthrowaway

                                                                                                                                                                                                                              today at 3:24 PM

                                                                                                                                                                                                                              Check out MTPLX and limit your context size.

                                                                                                                                                                                                                          • NorwegianDude

                                                                                                                                                                                                                            today at 3:16 PM

                                                                                                                                                                                                                            If the benchmarks are a real indication, we now have a local model that is runnable on a high-end personal PC that trades blows with the leading model Claude Opus 4.6 Max from half a year ago.

                                                                                                                                                                                                                            Insane if that is the case. Downloading now!

                                                                                                                                                                                                                              • throwaway613746

                                                                                                                                                                                                                                today at 5:54 PM

                                                                                                                                                                                                                                [dead]

                                                                                                                                                                                                                            • monkmartinez

                                                                                                                                                                                                                              today at 5:40 PM

                                                                                                                                                                                                                              Qwen3.6-27B has been the main LLM powering my little agentic stack. I have adopted the test and verify approach to any models allowed to run on my machine. When the "heretic" version drops, I will fire up the harness and test. Super excited to see how it stacks up against Qwen3.6!!!

                                                                                                                                                                                                                              • xlayn

                                                                                                                                                                                                                                today at 3:29 PM

                                                                                                                                                                                                                                The file "Just loads" on llama.cpp, the Unsloth https://huggingface.co/unsloth/Qwen3.8-27B-GGUF is an MTP file, I see mostly the same speed on pp and generation. There has to be something wrong with those benchmarks, I find extremely hard to believe a 27B model can work similar or exceed opus 4.6.

                                                                                                                                                                                                                                  • minimaltom

                                                                                                                                                                                                                                    today at 4:52 PM

                                                                                                                                                                                                                                    Worth distinguishing knowledge/task benchmarks from IF / agentic. It doesn't seem out of the question that you can have a small model thats generally good at instruction following and long-horizon agentic, as usually in those cases any requisite knowledge is in the context.

                                                                                                                                                                                                                                    Most of the benchmark improvements afaict are in agentic and instruction following benchmarks.

                                                                                                                                                                                                                                    • cyanydeez

                                                                                                                                                                                                                                      today at 5:59 PM

                                                                                                                                                                                                                                      I think you've been drinking the "LLMs only improve by adding parameter counts" that SOTA labs are selling VCs to build data centers so they can keep eating through cash to their own benefits.

                                                                                                                                                                                                                                      To the countrary, the reason Chinese models are excelling in the smaller area is because there's tons of fat in closed source models because of the crazy cash being thrown around.

                                                                                                                                                                                                                                      There absolutely is space to improve intelligence and capabilities without lathering on more and more parameters.

                                                                                                                                                                                                                                  • erdaltoprak

                                                                                                                                                                                                                                    today at 3:05 PM

                                                                                                                                                                                                                                    This is one of the most important model releases since most use cases don't need SOTA/Frontier

                                                                                                                                                                                                                                    If you want Qwen3.8-27B Serving Configs for the DGX Spark vLLM NVFP4 and RTX 4090 llama.cpp GGUF I added the setups here https://x.com/ErdalToprak/status/2088299678085308761?s=20

                                                                                                                                                                                                                                    • piyh

                                                                                                                                                                                                                                      today at 3:39 PM

                                                                                                                                                                                                                                      Qwen 3.6 is ~$2/m tok, 3.8 should be drop in replacement. Gemma 31B is $0.34/m tok. The price differential on these models is massive on openrouter.

                                                                                                                                                                                                                                        • SparkyMcUnicorn

                                                                                                                                                                                                                                          today at 5:19 PM

                                                                                                                                                                                                                                          Yeah, I would appreciate if someone could make sense of the pricing differences between these models. How can a provider run DSv4F at lower cost than a 27B dense or 35B A3B model?

                                                                                                                                                                                                                                          Does it come down to utilization and/or specific model tricks and efficiencies (attention, kv cache, etc.)?

                                                                                                                                                                                                                                          DeepInfra prices:

                                                                                                                                                                                                                                          Qwen 3.6 27B: $0.32 in / $3.20 out

                                                                                                                                                                                                                                          Gemma 3 27B: $0.08 in / $0.16 out

                                                                                                                                                                                                                                          DeepSeek V4 Flash 0731: $0.08 in / $0.18 out

                                                                                                                                                                                                                                          Qwen 3.6 35B A3B: $0.10 in / $0.95 out

                                                                                                                                                                                                                                          https://openrouter.ai/qwen/qwen3.6-27b

                                                                                                                                                                                                                                          https://openrouter.ai/google/gemma-3-27b-it

                                                                                                                                                                                                                                          https://openrouter.ai/qwen/qwen3.6-35b-a3b

                                                                                                                                                                                                                                          https://openrouter.ai/deepseek/deepseek-v4-flash-0731

                                                                                                                                                                                                                                            • mordae

                                                                                                                                                                                                                                              today at 5:58 PM

                                                                                                                                                                                                                                              DeepSeek V4 Flash is natively FP4 MoE with very compact KV cache. Say 8 GB/s. Qwen 27B is about 60 GB/s at full FP16 precision.

                                                                                                                                                                                                                                          • satvikpendem

                                                                                                                                                                                                                                            today at 6:03 PM

                                                                                                                                                                                                                                            Why are you comparing a 2.4 trillion Max model to a 31 billion model?

                                                                                                                                                                                                                                            • jjice

                                                                                                                                                                                                                                              today at 4:31 PM

                                                                                                                                                                                                                                              Where do you see that? From what I can see on Open Router, Qwen 3.6 27B (the closest dense equivalent to Gemma 31) is $0.28/m. Am I missing something?

                                                                                                                                                                                                                                              https://openrouter.ai/qwen/qwen3.6-27b

                                                                                                                                                                                                                                                • satvikpendem

                                                                                                                                                                                                                                                  today at 6:03 PM

                                                                                                                                                                                                                                                  They're comparing Qwen 3.8 Max to Gemma 31B, fundamental mistake.

                                                                                                                                                                                                                                                  • today at 4:40 PM

                                                                                                                                                                                                                                            • tosh

                                                                                                                                                                                                                                              today at 3:11 PM

                                                                                                                                                                                                                                              27b dense model at Opus 4.6 level

                                                                                                                                                                                                                                              Opus at home

                                                                                                                                                                                                                                              I hope there also will be a new ~10b variant

                                                                                                                                                                                                                                                • UncleOxidant

                                                                                                                                                                                                                                                  today at 4:08 PM

                                                                                                                                                                                                                                                  I'm hoping for a 3.8-122B MoE

                                                                                                                                                                                                                                                  • yassa9

                                                                                                                                                                                                                                                    today at 3:33 PM

                                                                                                                                                                                                                                                    can you tell me ideas of usecases of 9 or 10B language models ? I cant find any usecases other than training a lora on them to give good bash commands for example

                                                                                                                                                                                                                                                      • mring33621

                                                                                                                                                                                                                                                        today at 4:09 PM

                                                                                                                                                                                                                                                        9B Qwen models are good and fast for local python coding tasks.

                                                                                                                                                                                                                                                        • tosh

                                                                                                                                                                                                                                                          today at 3:49 PM

                                                                                                                                                                                                                                                          they are all overlapping but:

                                                                                                                                                                                                                                                          categorization, information retrieval, semantic search, image description

                                                                                                                                                                                                                                                          also with the model as part of an agentic system with tool calling

                                                                                                                                                                                                                                                          (edit: it is quite impressive what a small model in a feedback loop can do)

                                                                                                                                                                                                                                                  • syntaxing

                                                                                                                                                                                                                                                    today at 5:26 PM

                                                                                                                                                                                                                                                    Would I be surprised there’s bench maxing happening? Yes. But some users also use Q4 quantized and complain how dumb local models are.

                                                                                                                                                                                                                                                    • chvid

                                                                                                                                                                                                                                                      today at 3:08 PM

                                                                                                                                                                                                                                                      These are massive improvements - and something you can actually run on a laptop.

                                                                                                                                                                                                                                                      • mraza007

                                                                                                                                                                                                                                                        today at 4:33 PM

                                                                                                                                                                                                                                                        Man what a week, We just had GLM 5.3 that came out and then we had smaller local model Qwen3.8-27B from Qwen

                                                                                                                                                                                                                                                        Just tried using Pi Agent and looks very promising

                                                                                                                                                                                                                                                        • today at 4:48 PM

                                                                                                                                                                                                                                                          • mickeyp

                                                                                                                                                                                                                                                            today at 3:31 PM

                                                                                                                                                                                                                                                            Model benchmarks are useful, to a point, but it is the long tail of things you do with the model that determines if it's good at a wide range of activities. Ant/OAI, to their credit, build their models -- even the small ones -- so they follow instructions and do tool calling well, without the system prompts confusing them. This is especially important for long-horizon tool calling.

                                                                                                                                                                                                                                                            So one open weight model might "meet" Opus or whatever on benchmarks, but then fail to follow a simple answer format and also tool call correctly. The models are whipped to within an inch of their lives to strictly adhere to their post training quality gates.

                                                                                                                                                                                                                                                            • btbuildem

                                                                                                                                                                                                                                                              today at 5:57 PM

                                                                                                                                                                                                                                                              O joyous day!

                                                                                                                                                                                                                                                              • theanonymousone

                                                                                                                                                                                                                                                                today at 3:41 PM

                                                                                                                                                                                                                                                                I'm wondering whether any provider can offer this for cheaper $/token than the new DSv4 Flash, which is both cheaper and smarter :/

                                                                                                                                                                                                                                                                Completely local use is a different story, of course.

                                                                                                                                                                                                                                                                • minimaltom

                                                                                                                                                                                                                                                                  today at 4:10 PM

                                                                                                                                                                                                                                                                  Architecture thread! Afaict they continue to use gated attention + delta net, which was also adopted+adapted by K3, but im surprised theres no improvements to the residual stream (deepseek are using manifold hyper-connections, kimi have attention residuals) ?

                                                                                                                                                                                                                                                                  Perf improvements seem to all come from training?

                                                                                                                                                                                                                                                                    • anana_

                                                                                                                                                                                                                                                                      today at 4:19 PM

                                                                                                                                                                                                                                                                      As was the case with GLM 5.3, it seems that there is still much juice to be squeezed from post-training

                                                                                                                                                                                                                                                                  • bertili

                                                                                                                                                                                                                                                                    today at 4:20 PM

                                                                                                                                                                                                                                                                    Wow. Speed improved as well. 200t/s on a RTX 5090!

                                                                                                                                                                                                                                                                    https://x.com/sgl_project/status/2088281320422322413

                                                                                                                                                                                                                                                                    • arjie

                                                                                                                                                                                                                                                                      today at 4:16 PM

                                                                                                                                                                                                                                                                      I use the Qwens as a vision model for my DeepSeek V4 Flashes to handle. But the Qwens run on old RTX A6000 Ampere. Does anyone know if there's any news about INT4/AWQ quants for the RTX A6000?

                                                                                                                                                                                                                                                                        • ericd

                                                                                                                                                                                                                                                                          today at 4:26 PM

                                                                                                                                                                                                                                                                          Was recently thinking about doing something similar, do you basically just have the qwens describe what they see for the flashes?

                                                                                                                                                                                                                                                                          Was considering adding a LoRa/vision head to Flash, but seems like it could take a while to get it right.

                                                                                                                                                                                                                                                                          If DSv4 Flash was multimodal, I’d probably be done model shopping for a while

                                                                                                                                                                                                                                                                            • arjie

                                                                                                                                                                                                                                                                              today at 4:52 PM

                                                                                                                                                                                                                                                                              Same, with a multimodal DSv4 Flash I would just stop paying attention to things. Very smart, and at 260 tok/s it's too fast to care about anything else. If you ever graft something like that I would love to hear about it.

                                                                                                                                                                                                                                                                              Yes, I have a very dumb flow. The harness has a describe_image tool that takes an image and a prompt and so DSv4 Flash uses it to get an idea of what it's looking at.

                                                                                                                                                                                                                                                                                • ericd

                                                                                                                                                                                                                                                                                  today at 5:24 PM

                                                                                                                                                                                                                                                                                  Yeah, I might just replicate what you're doing. Main issue right now is just finding spare vram to actually run another model in parallel... And yeah, if I train up a vision adapter somehow, I'll try to put it up/post about it, seems like we're getting the killer apps for local LLMs right now, where it's just feasible enough if you're enthusiastic enough to be a bit economically irrational, and just useful enough to sort of rationalize.

                                                                                                                                                                                                                                                                      • TomGarden

                                                                                                                                                                                                                                                                        today at 3:02 PM

                                                                                                                                                                                                                                                                        Really excited to see what people do with this. 3.7 27B was probably the best compromise between size and intelligence to run on consumer hardware

                                                                                                                                                                                                                                                                        • synergy20

                                                                                                                                                                                                                                                                          today at 3:39 PM

                                                                                                                                                                                                                                                                          I wish this can run directly on my RTX 4090, seems like 30B is the sweet spot for dense model to run locally, sadly RTX 5090 is very expensive and I need a new PC and new power supply(and UPS) to run that, adding a second RTX 4090 is another option, but not sure if my PC can do that yet.

                                                                                                                                                                                                                                                                            • baron3dl

                                                                                                                                                                                                                                                                              today at 3:42 PM

                                                                                                                                                                                                                                                                              even a 3090 will give you the VRAM headroom. i run Q8 on an 3090/A6500 combo. well, Q8 of 3.6-27B. I'm building the Q8 GGUF for 3.8 now, assuming mine will finish before someone else's.

                                                                                                                                                                                                                                                                              • KyleJune

                                                                                                                                                                                                                                                                                today at 5:09 PM

                                                                                                                                                                                                                                                                                Others in this thread said it runs on RTX 4090.

                                                                                                                                                                                                                                                                            • ThouYS

                                                                                                                                                                                                                                                                              today at 3:17 PM

                                                                                                                                                                                                                                                                              I am so happy right now, qwen3.6-27b was an absolute game changer. To see another one in the same league.. phew

                                                                                                                                                                                                                                                                              • irthomasthomas

                                                                                                                                                                                                                                                                                today at 4:15 PM

                                                                                                                                                                                                                                                                                Why don't qwen/alibaba host the model themselves? I was looking forward to trying it on their coding plan. Google are the same way with their Gemma models.

                                                                                                                                                                                                                                                                                  • spwa4

                                                                                                                                                                                                                                                                                    today at 4:25 PM

                                                                                                                                                                                                                                                                                    Pretty sure you can use Gemma models on Google's "Vertex AI".

                                                                                                                                                                                                                                                                                • today at 3:10 PM

                                                                                                                                                                                                                                                                                  • kunver

                                                                                                                                                                                                                                                                                    today at 3:16 PM

                                                                                                                                                                                                                                                                                    Looks like a pretty significant improvement on the DeepSWE benchmark compared to the previous 27B model.

                                                                                                                                                                                                                                                                                    • kristopolous

                                                                                                                                                                                                                                                                                      today at 3:12 PM

                                                                                                                                                                                                                                                                                      q4km is about 48 tps on a 4090. my llama.cpp params are --flash-attn on --parallel 1 --load-mode mmap

                                                                                                                                                                                                                                                                                    • jlkivey

                                                                                                                                                                                                                                                                                      today at 3:50 PM

                                                                                                                                                                                                                                                                                      Note: on the model card the comparison to Opus is Opus 4.6 Max, not 4.7

                                                                                                                                                                                                                                                                                      • yassa9

                                                                                                                                                                                                                                                                                        today at 3:35 PM

                                                                                                                                                                                                                                                                                        Can anyone who has that specific personal test he tries on different models , and tries this model , to tell us here if possible , how good or bad is this new model ? compared to others ?

                                                                                                                                                                                                                                                                                        I only trust those users genuine personal tests

                                                                                                                                                                                                                                                                                          • alyandon

                                                                                                                                                                                                                                                                                            today at 3:39 PM

                                                                                                                                                                                                                                                                                            There is a down to earth guy on YT that performs a series of tests against LLMs running on non-god-tier commodity hardware. He will likely be testing this soon enough.

                                                                                                                                                                                                                                                                                            https://www.youtube.com/@lukesdevlab

                                                                                                                                                                                                                                                                                            I don't know if that is what you are looking for or not and as always your experiences may be different.

                                                                                                                                                                                                                                                                                              • yassa9

                                                                                                                                                                                                                                                                                                today at 4:14 PM

                                                                                                                                                                                                                                                                                                thaaanks man, this channel seems really informative, although < 10K subs only !

                                                                                                                                                                                                                                                                                                  • alyandon

                                                                                                                                                                                                                                                                                                    today at 4:20 PM

                                                                                                                                                                                                                                                                                                    It's a relatively new channel - but yeah - I feel the guy puts a lot of effort into what he does and deserves more subs.

                                                                                                                                                                                                                                                                                        • ThouYS

                                                                                                                                                                                                                                                                                          today at 3:30 PM

                                                                                                                                                                                                                                                                                          3.6-27B on little-coder was already mind blowing. looking forward to this guy!

                                                                                                                                                                                                                                                                                          • anana_

                                                                                                                                                                                                                                                                                            today at 3:15 PM

                                                                                                                                                                                                                                                                                            Monstrous benchmarks! Hoping it is not benchmaxxed.

                                                                                                                                                                                                                                                                                            • tosh

                                                                                                                                                                                                                                                                                              today at 3:21 PM

                                                                                                                                                                                                                                                                                              also cool: Qwen 3.8 27b is multi modal!

                                                                                                                                                                                                                                                                                                • gurkwart

                                                                                                                                                                                                                                                                                                  today at 3:42 PM

                                                                                                                                                                                                                                                                                                  strong visual reasoning apparently, which is nice. still lacking native audio however. hoping for more companies to embrace the spirit of something like `gemma-4-12b-qat` for actual multi-modality (text, image, video, audio).

                                                                                                                                                                                                                                                                                              • pu_pe

                                                                                                                                                                                                                                                                                                today at 3:25 PM

                                                                                                                                                                                                                                                                                                Seems to be SOTA for its size. Hopefully independent benchmarks will come soon.

                                                                                                                                                                                                                                                                                                • kunver

                                                                                                                                                                                                                                                                                                  today at 3:15 PM

                                                                                                                                                                                                                                                                                                  Welcome deepseek flash flash!

                                                                                                                                                                                                                                                                                                  • today at 3:15 PM

                                                                                                                                                                                                                                                                                                    • expedited123

                                                                                                                                                                                                                                                                                                      today at 3:14 PM

                                                                                                                                                                                                                                                                                                      Kinda was expecting to see Gemma 4 26B in benchmark comparisons :(

                                                                                                                                                                                                                                                                                                        • kamranjon

                                                                                                                                                                                                                                                                                                          today at 3:37 PM

                                                                                                                                                                                                                                                                                                          Since Qwen 3.6 27b outperforms Gemma 4 26b in most benchmarks I'm not sure the value - also Gemma 26b is a MOE model whereas this is a dense model, so not typically direct competitors at their sizes - Gemma 4 31b comparison would be interesting though.

                                                                                                                                                                                                                                                                                                            • expedited123

                                                                                                                                                                                                                                                                                                              today at 4:34 PM

                                                                                                                                                                                                                                                                                                              I see! Thanks.

                                                                                                                                                                                                                                                                                                      • davidw

                                                                                                                                                                                                                                                                                                        today at 5:19 PM

                                                                                                                                                                                                                                                                                                        I don't know much about the production of these models. How hard would it be to 'fork' something like this and have it not be full of CCP indoctrination?

                                                                                                                                                                                                                                                                                                          • regularfry

                                                                                                                                                                                                                                                                                                            today at 5:43 PM

                                                                                                                                                                                                                                                                                                            Look for `heretic` fine-tunes in the next couple of days.

                                                                                                                                                                                                                                                                                                        • lossolo

                                                                                                                                                                                                                                                                                                          today at 5:09 PM

                                                                                                                                                                                                                                                                                                          Why weren't the points merged again from the "dupe" thread that had 289 points?

                                                                                                                                                                                                                                                                                                          https://news.ycombinator.com/item?id=49299684

                                                                                                                                                                                                                                                                                                          What a weird mechanism. If someone is judging a thread/topic/event impact by the number of points it got, then doing this unfairly degrades that thread.

                                                                                                                                                                                                                                                                                                          It should have deduped by user and combined the 168(at the time of writing this comment) + 289 points. Just add the twitter link from the previous thread as an additional link in the description, like you normally do, move all the points over, and remove the old thread.

                                                                                                                                                                                                                                                                                                          • naasking

                                                                                                                                                                                                                                                                                                            today at 5:07 PM

                                                                                                                                                                                                                                                                                                            Can anyone confirm whether this new Qwen release is any more concise when thinking? Overthinking was the biggest (only?) downside of the Qwen models.

                                                                                                                                                                                                                                                                                                            • altruios

                                                                                                                                                                                                                                                                                                              today at 3:14 PM

                                                                                                                                                                                                                                                                                                              remember to let llama.cpp catch up to anything new in this model. Save your judgment until about 2 weeks of use.

                                                                                                                                                                                                                                                                                                                • chrismartin

                                                                                                                                                                                                                                                                                                                  today at 3:33 PM

                                                                                                                                                                                                                                                                                                                  'Good' news, there seems to be nothing new architecture-wise. Same as Qwen 3.5 and 3.6, so llama.cpp doesn't know the difference.

                                                                                                                                                                                                                                                                                                              • filup

                                                                                                                                                                                                                                                                                                                today at 3:47 PM

                                                                                                                                                                                                                                                                                                                https://news.ycombinator.com/item?id=48403639

                                                                                                                                                                                                                                                                                                                my prediction was way too far out. 4.6 at home! Woo.

                                                                                                                                                                                                                                                                                                                • tristor

                                                                                                                                                                                                                                                                                                                  today at 4:59 PM

                                                                                                                                                                                                                                                                                                                  I'm hoping to see folks distill this with current generation Opus / Fable reasoning traces. I have had my best results locally so far from Qwopus (Qwen 3.6-27B w/ Opus 4.6 reasoning distilled). This looks GREAT and I am definitely setting this up later today.

                                                                                                                                                                                                                                                                                                                  • brcmthrowaway

                                                                                                                                                                                                                                                                                                                    today at 3:25 PM

                                                                                                                                                                                                                                                                                                                    This with ddg mcp to fill in world knowledge. Are local models the future when computer architectures catch up?

                                                                                                                                                                                                                                                                                                                    • alpha_trion

                                                                                                                                                                                                                                                                                                                      today at 3:16 PM

                                                                                                                                                                                                                                                                                                                      NICE, i've been waiting for this drop, thanks for posting this

                                                                                                                                                                                                                                                                                                                      • Mr_Eri_Atlov

                                                                                                                                                                                                                                                                                                                        today at 3:38 PM

                                                                                                                                                                                                                                                                                                                        This is the homelab model hands down

                                                                                                                                                                                                                                                                                                                        • brcmthrowaway

                                                                                                                                                                                                                                                                                                                          today at 3:11 PM

                                                                                                                                                                                                                                                                                                                          My Strix Halo is about to go overdrive!

                                                                                                                                                                                                                                                                                                                          • cmrdporcupine

                                                                                                                                                                                                                                                                                                                            today at 4:46 PM

                                                                                                                                                                                                                                                                                                                            I found this kind of amusing while running it (using Pi as the harness). Don't know if this is evidence of intense fine tuning from Claude but it smells like it...

                                                                                                                                                                                                                                                                                                                            " The user wants me to explore the repository at XXXX and report back. Let me start by understanding the project structure, reading the CLAUDE.md file, and getting a general overview of what this repository is.

                                                                                                                                                                                                                                                                                                                            Let me start by reading the main project documentation and exploring the directory structure.

                                                                                                                                                                                                                                                                                                                            I'll take a look around this repo. Let me start by getting a lay of the land.

                                                                                                                                                                                                                                                                                                                            read resource CLAUDE.md (ctrl+o to expand)

                                                                                                                                                                                                                                                                                                                            ENOENT: no such file or directory, access 'XXXX/CLAUDE.md'"

                                                                                                                                                                                                                                                                                                                            • ramon156

                                                                                                                                                                                                                                                                                                                              today at 3:28 PM

                                                                                                                                                                                                                                                                                                                              need another fable uncensored merge with 3.8, really curious what it can deliver

                                                                                                                                                                                                                                                                                                                              • ggerganov

                                                                                                                                                                                                                                                                                                                                today at 6:02 PM

                                                                                                                                                                                                                                                                                                                                [dead]

                                                                                                                                                                                                                                                                                                                                • today at 5:44 PM

                                                                                                                                                                                                                                                                                                                                  • RobertasTa

                                                                                                                                                                                                                                                                                                                                    today at 4:24 PM

                                                                                                                                                                                                                                                                                                                                    [flagged]

                                                                                                                                                                                                                                                                                                                                    • fintuner

                                                                                                                                                                                                                                                                                                                                      today at 4:13 PM

                                                                                                                                                                                                                                                                                                                                      [flagged]

                                                                                                                                                                                                                                                                                                                                      • steffi_oliver

                                                                                                                                                                                                                                                                                                                                        today at 4:50 PM

                                                                                                                                                                                                                                                                                                                                        [flagged]

                                                                                                                                                                                                                                                                                                                                        • today at 3:29 PM