\

Xiaomi Mimo 2.6 live post-training dashboard

419 points - yesterday at 8:09 PM

Source
  • joelwallis

    yesterday at 8:54 PM

    I been using MiMo-V2.5 to do most of my work as software engineer, on a variety of projects I'm working on, and I been VERY happy with ROI. The model is very powerful! Not perfect – I've run in hallucination loops once or twice, but nothing a stop-then-continue wouldn't solve.

    The cost is unbelievably low, and the quality of intelligence I get is equivalent to when I was working mostly with Anthropic models (late last year/early this year). I'm fully invested in MiMo and I'm very happy with it.

    -- PS: I also check almost daily to see if other models are capable of doing such great work. And they do – DS4F is powerful and DS41 is impressive, GLM 5.3 Flash gets a job done well, etc. – but when I add cost of M-token in the ROI math, Jeez! MiMo is an order of magnitude better.

      • miyuru

        today at 5:59 AM

        Same here. It’s the first AI provider I actually gave money to, since they offered the model for free with a Mimo code for the first month or so, and it was great.

        These days, there are more intelligent models like DS4.1, but Mimo is very obedient, so I plan things with another model and give the implementation to Mimo.

        • baxtr

          today at 6:31 AM

          Could you elaborate on how you check daily? Do you swap models for certain tasks?

          • alwinaugustin

            yesterday at 11:21 PM

            I am also using 2.5 and it is giving me solid results. Its available free on Openrouter

            • rapind

              today at 1:25 AM

              I’ve been very pleased with DS 4.1 flash. Not so much the 4.0 models, but for coding (Rust) it’s been great so far (3 solid days of work).

              I’ll give Mimo a try.

                • trollbridge

                  today at 1:47 AM

                  MiMo is my backup whenever DeepSeek is down, had the price bump, is slow, etc.

                  UltraSpeed was absolutely awesome. I miss it.

                  DS 4.1 Flash is amazing. Well worth the extra cost.

              • walrus01

                yesterday at 9:15 PM

                I've found that mimo v2.5 works for very basic things like a python script to do one thing, but it also is very 'dumb' compared to qwen 3.8-flash-next (I think the benchmark scores for terminal and coding specific benches back this up). And definitely not in the same class as like a GLM5.2 or 5.3. It's fast but makes basic mistakes that only get caught later.

                  • girvo

                    today at 2:19 AM

                    The fact I can run Qwen 3.8 Flash Next locally, forever (on my DGX Spark-alike) is genuinely shocking to me. It’s crazy good for how small it is. Fast, too.

                      • jonsoft

                        today at 4:27 AM

                        I made this 3D game in a day on the same setup with Qwen Code as agent: https://games.jonathanpage.com/

                        And I am not a web developer! It's an extraordinary model.

                        (Mouse and keyboard required)

                        • walrus01

                          today at 3:01 AM

                          Yeah, I'm guessing you have a variant that fits in <128GB with 262k context? I have the unsloth Q8 GGUF of it here in a setup that with full context and ton of extra llama-server "--cache-ram" sits around 200GB RAM usage on a 256GB system, it's probably the best thing I've found for a 256GB class machine. Enough headroom for a rope/yarn extension to 524288 context if I need it.

                            • girvo

                              today at 3:57 AM

                              Yep, the engrams are on NVMe (the speed penalty was lower than I expected) and it is quantised to fit.

                              It’s good enough that I’m considering a second spark, or selling this and buying an M5 Ultra with 256GB for it

                      • yesterday at 10:58 PM

                    • flexagoon

                      yesterday at 11:45 PM

                      How does it compare with DS 4.1 Flash in your experience, if you ignore the cost?

                      • jwpapi

                        yesterday at 9:37 PM

                        May I ask why you ended up there instead of just using the heavy subsidized subscription. I’m actually curious.

                          • eli

                            yesterday at 11:18 PM

                            Mimo has subsidized subscriptions too

                        • james2doyle

                          yesterday at 8:58 PM

                          2.5 Pro or the regular 2.5?

                          I always found that those Mimo models to be really good at tool calling and following instructions

                          • ignoramous

                            today at 6:46 AM

                            > GLM 5.3 Flash gets a job done well, etc. – but when I add cost of M-token in the ROI math, Jeez! MiMo is an order of magnitude better.

                            API may be expensive, but I do 900m tokens (95% cached, ~0.4% output) on Z.ai's $18/mo coding plan with GLM 5.3 Flash.

                              • miroljub

                                today at 7:04 AM

                                I wouldn't call that inexpensive.

                                For comparison, I am currently at 6.6B tokens, 95% of monthly quota on a 10$ command code plan, mostly using DeepSeek flash 4.1, or some of the free models for easier tasks.

                            • esafak

                              yesterday at 9:55 PM

                              How fast is it compared with the other Chinese models?

                                • ricardobeat

                                  yesterday at 10:59 PM

                                  They both are in the 50-100 tok/s range. The Mimo v2.5 Pro Ultraspeed beta could reach 1000 tok/s, hoping they can do something similar for the new model, it was amazing.

                              • wangxili1997

                                today at 8:05 AM

                                [flagged]

                                • electroglyph

                                  yesterday at 11:50 PM

                                  [flagged]

                                    • NuclearPM

                                      today at 12:01 AM

                                      Real?

                                        • electroglyph

                                          today at 2:59 AM

                                          mimo 2.5 has been a big underperformer since shortly after it's release imo. i cancelled my sub after the first month. purposefully using 2.5 right now is just handicapping yourself for no reason.

                                            • NuclearPM

                                              today at 3:22 AM

                                              I understand now. You used the wrong word.

                                  • yeeeloit

                                    yesterday at 9:38 PM

                                    [flagged]

                                      • senordevnyc

                                        yesterday at 9:51 PM

                                        Yeah, this Brazilian dude who has been a contributor here on HN longer than your anonymous account is shilling for a Chinese model company. Makes sense.

                                        • platinumrad

                                          yesterday at 9:49 PM

                                          Are you accusing them of astroturfing? Why is it strange for someone to say something topical?

                                  • dr_dshiv

                                    yesterday at 10:37 PM

                                    Well, if open source AI is dangerous (for OpenAI/Anthropic IPOs?), this is like watching a time bomb.

                                      • dzonga

                                        today at 12:00 AM

                                        the open burial started when zAI served their latest model on all Chinese chips.

                                        now we r just noticing the grave getting dug deeper.

                                        • skybrian

                                          today at 1:00 AM

                                          For my own usage, Luna is cheap enough that I don't care if other models are cheaper. I'm interested if another model is in some way better and not too expensive.

                                            • rapind

                                              today at 1:28 AM

                                              Luna is great but makes a lot of mistakes at high and lower in my experience (large rust codebase). I use Luna Max for asynchronous subagent reviews and am very happy with its work, but it’s slow af.

                                              • ijidak

                                                today at 2:52 AM

                                                What plan are you on?

                                                Trying to understand why users are using Luna when Sol seems essentially unlimited on the pro plan. Unless you have jobs running 24/7.

                                                  • teki_one

                                                    today at 3:15 AM

                                                    Sol is useless atm on the Plus plan, 1-2 questions 5-10m to get through the 5h allowance. (used to be good, can change any day)

                                            • SlightlyLeftPad

                                              today at 1:11 AM

                                              I think it might be closer to this:

                                              https://www.debtdefaultclock.us/

                                          • passive

                                            yesterday at 10:46 PM

                                            Neat! I've been trying out their next model for the last week, which I assume is a version of this, and it's been a good experience so far.

                                            I had used 2.5-pro for a hefty chunk of development, and found it to work like a somewhat forgetful senior engineer who was new to my project. Very capable, would almost always choose a reasonable option, if not always the best one for the project, and not great at multi-tasking. Generally, made me comfortable not scrutinizing the code line-by-line, but still needed a bit of steering once projects got to a reasonable size.

                                            The next model is a clear step up in the multi-tasking capability at least, with me very rarely having to steer the implementation of a well-defined issue. In terms of code, I found MiMo-V.2.5-pro to be extremely conservative, implementing minimal solutions. The next model seems a little bit more ambitious, in positive ways, making good guesses about gaps/next steps. It also seems to be a fair bit better at design, at least for the little bit I've done, it was good at translating my concepts to practical elements on screen, and cleaned things up nicely as I made suggestions.

                                            • ricardobeat

                                              yesterday at 10:55 PM

                                              For reference, Mimo-v2.5-Pro scored 19% on DeepSWE 1.1. This is looking great.

                                              Fable scores 70%, Kimi K3 69%, Astra 74% (all on max effort).

                                              https://deepswe.datacurve.ai/blog/deepswe-v1-1

                                                • Cookingboy

                                                  today at 12:18 AM

                                                  2.6-pro just reached 63.7% by step 10, it's on step 11 right now.

                                                  Even flash reached 60.7% by step 12, and it's on step 16 now.

                                                  This is so exciting lmao.

                                              • krm01

                                                yesterday at 8:44 PM

                                                This is pretty neat. What would be a good reason for the other Model providers to not do this?

                                                  • kibae

                                                    yesterday at 8:54 PM

                                                    Speculating here, but I assume researchers can make a reasonable estimate of the size of closed models based on factors like training time, training speed, and the number of tokens processed.

                                                    Also, Anthropic and OpenAI probably want to keep each other on their toes so they don’t end up on the wrong side of another Opus 4.6 / GPT-5.3-Codex situation, where one lab releases a model only for the other to drop a better one hours later.

                                                      • Bolwin

                                                        today at 5:57 AM

                                                        I don't really remember a situation, which of those models supposedly beat the other?

                                                        I still opus 4.6 though not for code

                                                        • jwpapi

                                                          yesterday at 9:35 PM

                                                          I think first of all it’s not an obvious idea, also the marketing surplus for other providers is not as big for openai/anthropic as for xiaomi and last but not least I’m pretty sure you can withdraw methodology from here.

                                                          I’m saying who has a million dollars for me, so I can make my own model?

                                                      • nikcub

                                                        today at 3:49 AM

                                                        this is remarkable transparency in an otherwise hyper competitive and secretive industry

                                                    • liuliu

                                                      yesterday at 8:52 PM

                                                      When you run benchmarks while training, isn't that the definition of contamination? Asking because I am not sure if this is normal in big labs now.

                                                        • jampekka

                                                          yesterday at 9:00 PM

                                                          Kinda yes. The benchmarks become part of the validation set, which means the models get slightly overfit to them if they are used as criteria for stopping the training. But a lot less compared to using them in the training data.

                                                          I'd guess everybody uses at least some benchmarks as stopping criteria, which is kinda sensible, but it also does induce some benchmaxxing, and explains partly why the newest models always tend to eke out in benchmarks.

                                                          https://en.wikipedia.org/wiki/Training,_validation,_and_test...

                                                            • liuliu

                                                              yesterday at 9:04 PM

                                                              Correct. If just stopping criteria, that is less contaminated. The question gets muddier once you also use it to determine hyperparameters during small-scale runs.

                                                          • lucrbvi

                                                            yesterday at 8:55 PM

                                                            They are using it to evaluate checkpoints during the training, they are probably not using the benchmarks for training the models. It's a common practice for big reinforcement learning runs.

                                                            • nodja

                                                              yesterday at 10:19 PM

                                                              They exist to detect degradation. Datasets are not perfect and if a batch contains too much bad data it can ruin a run, also an opportunity to find bad data and improve the dataset filtering.

                                                              • SwellJoe

                                                                yesterday at 9:02 PM

                                                                You gotta have something to aim at. And, presumably, the benchmark is not part of the training data, it is the test against which the model is tested at each stage; is behavior moving in the right direction?

                                                                • esafak

                                                                  yesterday at 9:57 PM

                                                                  Not if you don't train against them.

                                                                    • kingstnap

                                                                      yesterday at 10:27 PM

                                                                      It's implicitly trained against. There is like information leakage with researchers messing with the training parameters and checkpoints used.

                                                                      It's not the direct feedback loop of RL but its not far.

                                                              • fzysingularity

                                                                yesterday at 10:15 PM

                                                                Very cool to see the openness here, and likely more like this will come from smaller startups where they win users on transparency.

                                                                • ssn2000

                                                                  today at 4:26 AM

                                                                  Total run cost is $1.2M until now, what resources are they using to train their model? Wish they shared more details on that and what the MFU metrics are.

                                                                  • ProfessorLayton

                                                                    yesterday at 8:52 PM

                                                                    2.6 Pro: >started 2026-09-15 10:32 UTC

                                                                    For some reason I thought training took much, much longer than what the progress bar suggests.

                                                                    This is really neat, I'm currently using mimo 2.5 pro, and it's decent (or great given the price). Hopefully their next one is multimodal.

                                                                      • GaggiX

                                                                        yesterday at 8:58 PM

                                                                        These are post-training reinforcement learning steps.

                                                                          • krackers

                                                                            yesterday at 9:00 PM

                                                                            Yes, updated the submission title to say "post-training" to hopefully prevent further confusion

                                                                    • ttul

                                                                      today at 1:33 AM

                                                                      $5 per second if my eyes don’t fool me. That’s ~$432K per day. Enough to rent 3,000 B300 nodes on Modal.

                                                                      • thehamkercat

                                                                        yesterday at 8:51 PM

                                                                        This is crazy, but sadly anthropic/openai will never do this, what has happened to this world, where chinese companies are more open than US or even EU companies

                                                                          • medlazik

                                                                            yesterday at 9:18 PM

                                                                            Neoliberalism, that famously open and transparent economic ideology

                                                                              • atemerev

                                                                                today at 5:08 AM

                                                                                Ah, one Donald Trump, a famous neoliberal.

                                                                        • speedgoose

                                                                          yesterday at 8:47 PM

                                                                          I didn't know 2 thirds of the training data would be source code.

                                                                            • jerrygenser

                                                                              yesterday at 8:53 PM

                                                                              that is the the "data used to improve the model" when signing up for the subscription plans

                                                                              • leothetechguy

                                                                                yesterday at 8:53 PM

                                                                                this is the rl run, not the pretraining run

                                                                                  • ahmadyan

                                                                                    yesterday at 9:28 PM

                                                                                    even in pre-training, usually 30%-50% is code these days.

                                                                            • kkotak

                                                                              today at 3:26 AM

                                                                              Wouldn't us observing this break down the model superposition and make it dumber? :)

                                                                                • brcmthrowaway

                                                                                  today at 3:41 AM

                                                                                  Found the Dark Matter (2024) watcher

                                                                              • jstummbillig

                                                                                today at 5:26 AM

                                                                                Wow, spending money on training an almost-frontier-model is much more time intensive than I thought it was.

                                                                                • rao-v

                                                                                  today at 2:00 AM

                                                                                  I absolutely love that someone is doing this! Why isn’t IBM for Granite or Google for Gemini?

                                                                                  If you are going to develop a near frontier model, and you don’t think you have special sauce up your sleeve, why not making training runs and RL environment scores etc. visible to the world?

                                                                                  I’m genuinely learning quite a bit just from the dashboard

                                                                                    • gtirloni

                                                                                      today at 2:44 AM

                                                                                      They think they have the special sauce. Even if they do, what would they get in return for doing that?

                                                                                  • wg0

                                                                                    today at 6:33 AM

                                                                                    "Slow down this much openness in AI or we won't get our trillion dollars valuations!"

                                                                                    Google had this GPT long go and a wise man within Google noted:

                                                                                    "We don't have any maot neither does anyone else."

                                                                                    The AI bubble burst is guaranteed and is only delayed by IPOs.

                                                                                    • rozab

                                                                                      yesterday at 8:52 PM

                                                                                      Why are they doing this? To try head off accusations about distillation?

                                                                                        • bayindirh

                                                                                          yesterday at 8:59 PM

                                                                                          Sometimes you're confident about what you're doing and show how you work to the world.

                                                                                          Keeping the garage door open, or at least making the door translucent. It's always cool.

                                                                                          • jampekka

                                                                                            yesterday at 9:15 PM

                                                                                            That China's official policy is now to prefer open models and open model development may be a part of it.

                                                                                              • culi

                                                                                                yesterday at 9:39 PM

                                                                                                BRICS just had a New Delhi meeting where Xi pushed a 5-point plan on AI cooperation that centered on open source models

                                                                                                • Aboutplants

                                                                                                  yesterday at 9:21 PM

                                                                                                  With that policy in place, labs might be incentivized to be creative in their openness. This being fun/free PR

                                                                                              • anemic

                                                                                                yesterday at 9:53 PM

                                                                                                Bottom of the page says "Open is what we value."

                                                                                            • wolttam

                                                                                              yesterday at 8:42 PM

                                                                                              Hah, it would be great to see more labs pick this up.

                                                                                              • ernsheong

                                                                                                yesterday at 11:03 PM

                                                                                                Mino 2.5 has been my workhorse for coder and tester agents (the ones planner agents delegate tasks to)

                                                                                                • thenews

                                                                                                  today at 2:29 AM

                                                                                                  been using the 2.5 mimo for side projects, works amazing

                                                                                                  • dr_kiszonka

                                                                                                    today at 1:17 AM

                                                                                                    Very curious that everyone here (so far) seems to assume this dashboard presents real data.

                                                                                                      • hsbalanxvxjsmab

                                                                                                        today at 2:39 AM

                                                                                                        Haha yeah pretty wild how easily you can see the data is fake by the repeating numbers (refresh the page the progress goes back in time constantly) + watch for restarts. They say they happen but 0 data correlates the log messages. Just a replay of old data or being fed by an llm so they convince people they are open

                                                                                                          • Bolwin

                                                                                                            today at 6:00 AM

                                                                                                            The intermediate tickers are fake but real data comes in and resets it. Its like a progress bar essentially. We don't call progress and bars fake

                                                                                                    • esafak

                                                                                                      yesterday at 9:53 PM

                                                                                                      That's the kind of transparency we need! That DeepSWE benchmark puts it in frontier territory: https://artificialanalysis.ai/agents/coding-agents?coding-ag...

                                                                                                      • Alifatisk

                                                                                                        today at 7:03 AM

                                                                                                        Can we call this open AI?

                                                                                                        • tcbbd

                                                                                                          today at 3:04 AM

                                                                                                          [dead]

                                                                                                          • heronbank

                                                                                                            today at 1:35 AM

                                                                                                            [dead]

                                                                                                            • sinuhe69

                                                                                                              today at 6:27 AM

                                                                                                              [flagged]

                                                                                                              • Toslink

                                                                                                                yesterday at 11:04 PM

                                                                                                                [dead]

                                                                                                                • pppkin

                                                                                                                  today at 2:33 AM

                                                                                                                  [dead]

                                                                                                                  • impulser_

                                                                                                                    yesterday at 9:22 PM

                                                                                                                    The Chinese labs are just making fun of the US labs at this point.

                                                                                                                    Where is the cool shit from the US labs?

                                                                                                                      • culi

                                                                                                                        yesterday at 9:41 PM

                                                                                                                        With other software, devs convince their managers of the importance of using open source stuff in their stack. With AI, it's usually managers choosing what models to use for the devs. The US labs don't need to give a damn how much devs like open source

                                                                                                                          • impulser_

                                                                                                                            yesterday at 10:45 PM

                                                                                                                            This isn't about liking open source. This is about the labs just being cool and doing cool shit instead of the opposite which is Anthropic where all they talking about is killing everyone and taking everyone's job.

                                                                                                                              • dlisboa

                                                                                                                                today at 1:25 AM

                                                                                                                                These labs are still (for the time being) made of people, who reflect their lives onto the work.

                                                                                                                                The US population is much more pessimistic and doomsday driven these days, whereas the Chinese are more optimistic and future driven.

                                                                                                                            • noir_lord

                                                                                                                              yesterday at 10:18 PM

                                                                                                                              > The US labs don't need to give a damn how much devs like open source

                                                                                                                              In the short term, true.

                                                                                                                              In the long term, unknown but typically when you hold progress that way while other countries don't you at best end up becoming siloed while the rest of the world continues on without you.

                                                                                                                          • hsbalanxvxjsmab

                                                                                                                            today at 2:42 AM

                                                                                                                            You mean all of the frontier models that the Chinese distillation clones are copying? Yeah kinda cool imo. If a dashboard showing training for a model that doesn't even come close to anything us labs have released in 6 months is "cool", then you're a loser

                                                                                                                              • bicepjai

                                                                                                                                today at 3:13 AM

                                                                                                                                Hahaha. Is that Sam or Dario with throwaway account. This sounds like calling social security, a free handout. Who distills the distillaters? Get it?

                                                                                                                                • impulser_

                                                                                                                                  today at 5:05 AM

                                                                                                                                  Why the fuck would you or I care about that?

                                                                                                                                  Anthropic and OpenAI literally stole from every human in history and youre out here complaining that the Chinese are distilling models and releasing them to the public?

                                                                                                                                  Why do you care?

                                                                                                                                  • atemerev

                                                                                                                                    today at 5:09 AM

                                                                                                                                    No crying in the copyright casino.

                                                                                                                            • dude250711

                                                                                                                              yesterday at 11:26 PM

                                                                                                                              Distillation in real-time? Very interesting!

                                                                                                                                • Cookingboy

                                                                                                                                  today at 12:30 AM

                                                                                                                                  That "training cost" is just live revenue count for Anthropic/OpenAI API calls!

                                                                                                                                  /s

                                                                                                                              • hsbalanxvxjsmab

                                                                                                                                today at 2:37 AM

                                                                                                                                This is so very clearly fake? See the message stating the flash 2.6 flash run was restarted and 0 graphs correlate that restart

                                                                                                                                  • Retro_Dev

                                                                                                                                    today at 4:19 AM

                                                                                                                                    A restart of the process does not necessarily mean reverting the model state. I don't know why you would even do that, because you'd lose all the progress you made.

                                                                                                                                • levocardia

                                                                                                                                  yesterday at 8:57 PM

                                                                                                                                  You'd think they would make it less obvious that they are running their whole operation with Claude

                                                                                                                                    • ricardobeat

                                                                                                                                      yesterday at 11:01 PM

                                                                                                                                      If you're thinking of the UI style, definitely not Claude. It is incapable of writing a clear sentence like "what each step's samples are made of", would have used all-caps for everything, more padding and gradients.

                                                                                                                                        • conception

                                                                                                                                          today at 2:26 AM

                                                                                                                                          I hope this is /s because it’s very easy to get Claude to write sensibly. That’s why AI slop writing is so annoying because it’s so easy to avoid with any amount of effort at all.

                                                                                                                                      • iammrpayments

                                                                                                                                        today at 6:32 AM

                                                                                                                                        Did you come here to astroturf or are you a big fan of Claude

                                                                                                                                        • jambutters

                                                                                                                                          today at 12:06 AM

                                                                                                                                          They'd be running in the red then cause they charge way less than Claude. Sorry but it just doesn't make logical sense. They have open source, papers, and self hosting too

                                                                                                                                          • SwellJoe

                                                                                                                                            yesterday at 9:00 PM

                                                                                                                                            It's not obvious to me. What's the tell?

                                                                                                                                            • cpcabbge

                                                                                                                                              today at 6:29 AM

                                                                                                                                              Do tell cause I can't