\

K2 Horizon: A connected fleet of six open models

277 points - yesterday at 3:36 PM

Source
  • jjordan

    yesterday at 4:22 PM

    Fully open models really need to be a big part of the AI future. That includes all source code, open training data, how it's organized, fed to the model, processed, etc. Until that becomes a thing you're always going to be left wondering what exactly lies underneath the closed model you are using, leaving open the possibility for societal manipulation.

      • culi

        yesterday at 10:04 PM

        Other than open training data (currently legally impossible), all of this holds for basically every major Chinese-made model. They not only open the weights but publish detailed methodology papers alongside the models in arXiv and even open source the code.

        • kibae

          yesterday at 4:49 PM

          The training data would need to have a permissive license for this to be possible.

            • embedding-shape

              yesterday at 6:26 PM

              Or, we just need to get this over with and declare any digital data findable via the internet to just be public property of everyone. Everything becomes public, besides stuff you keep locally, and there is no difference anymore, it's all just data anyone can use for whatever. A 1 year grace period for everyone to pull stuff off they don't want to be a part of this bright new open era, then we just scrap everything related to intellectual property, copyright and similar stupid stuff, and slap UBI on top of all of it for good measure.

                • chme

                  yesterday at 7:08 PM

                  I'd prefer to stay within the [hacker ethics](https://www.ccc.de/en/hackerethics), and protect private data. For non-private/personal data, sure. But individual people need their privacy protected.

                    • embedding-shape

                      yesterday at 9:09 PM

                      Me too, I'm hacker ethics all the way, which is why I'm saying anything network connected should really realize the "All information should be free." dream, and then private data should be far away from the internet, on computers/drives not even connected to the internet. The whole E2E encryption is a ticking time bomb people rely to keep their data safe from others, but nothing that you don't physically have close to you can be truly secret forever, and even then it'll be hard.

                  • jchw

                    today at 2:59 AM

                    I hate to invoke Poe's law but, I've now flip-flopped like six times over whether this could be serious.

                    I think it is serious. In which case, I gotta say, it really seems like you didn't spend much time thinking about this. "A 1 year grave period for everyone to pull stuff off they don't want to be a part of" - How does that work when the Internet is already full of unauthorized reproductions, most of which people aren't even aware of? Even ignoring practical considerations, when literally everyone is basically stuck using the Internet for everything, this seems a bit unfair to anyone who isn't onboard, akin to The Onion's Google Opt-out Village. But there are so many practical issues with this, it would be easier to list the number of problems this doesn't have. You accidentally leak something to the Internet and it becomes commons? What happens when other people leak things to the Internet? How about revenge porn?

                    Not minor stuff that can easily be papered over, this literally reintroduces the problem of needing to care about the provenance of data again, in a way that can't be automated, which makes the whole thing entirely moot. All just to make training data for AI models easier to distribute?

                    I'm all for intellectual property reform, maybe even fairly radical. But this just seems like it wasn't thought out.

                    If this was satire, well, I took the bait. Oddly convincing despite being hard to believe.

                    • tshaddox

                      yesterday at 8:44 PM

                      I don't really get what you're suggesting. You give a 1 year grace period for Metallica to pull all its music off the Internet, but then as soon as I host some of their MP3s on my Wordpress blog it's "public property of everyone" from that point forward?

                      • yesterday at 7:03 PM

                        • jrm4

                          yesterday at 8:43 PM

                          What you're slightly more realistically looking for here is for publicly available data to have a Fair Use exemption for certain uses, which is certainly something worth discussing.

                          • dotancohen

                            yesterday at 7:51 PM

                            I respectful disagree. I enjoy reading e.g. Asimov and well-executed journalism. And I completely respect the IP of those people who create these works.

                              • idiotsecant

                                today at 12:28 AM

                                I'm what way does it make sense to respect the intellectual property rights of a dead man?

                                  • ceejayoz

                                    today at 1:42 AM

                                    If IP rights end at death, there are some significant perverse incentives for offing big-name artists and authors and whatnot.

                                      • rmunn

                                        today at 3:06 AM

                                        The current "lifetime of the author + X years" rules in effect in the United States still carry the same perverse incentive, though the incentive diminishes rapidly as X gets larger; with the very large value of X in effect today the perverse incentive is so small as to be effectively non-existent, but it's still there in theory.

                                        Personally, I'd prefer a fixed term. I know enough independent authors making a living from selling their books that I'm willing to allow the fixed term to be large, like 50 years from date of completion of the work. (With a good definition of "completion" so someone can't cheat by editing a couple lines per year to keep something copyrighted indefinitely). The simpler the rule is, the easier it is to understand, and the harder it is to cheat it. The more complicated you make a rule, the more loopholes get found.

                                        • idiotsecant

                                          today at 4:03 AM

                                          There are significant perverse incentives for me shooting you with a gun and taking your money and running away too. It mostly doesn't happen.

                                      • dotancohen

                                        today at 12:39 AM

                                        That dead man took the risk of not earning much in his lifetime, to continue feeding his family even after his passing. You might as well ask why does someone acquire life insurance.

                            • ux266478

                              yesterday at 4:59 PM

                              You could sidestep it by running non-permissibly licensed training data that you purchased through an LLM. Legal attitude so far seems to be that this is transformative as long as it's not 1:1. The question on whether or not the end result is copyrightable of course remains controversial and inconsistent, but that question is also fairly irrelevent. You don't get more libre than public domain.

                              That's a fair amount of computational and labor overhead mind you, as you'll need to verify and prune the quality of your mountain of synthetic data, but certainly possible.

                              Though this assumes the legal system is a rational actor playing by the set of rules it claims to. In fact, I highly suspect you could get very unlucky and get an unfavorable ruling against you, because you stepped on a big pile of money's toes in the process of doing this.

                                • alightsoul

                                  yesterday at 6:29 PM

                                  It can also be used to sidestep copyright like this forum, books and most websites even if the data was not purchased but is a website or book.

                                  Are LLMs what we need to make all data public domain? This way it could be used for that purpose

                              • jjordan

                                yesterday at 6:10 PM

                                Hear me out.

                                Decentralized unstoppable storage, combined with decentralized unstoppable training, sorta like SETI for AI training. The seed of this tech already exists with IPFS and others like it.

                                We know (some? all?) of the big labs have skirted copyright laws at one point or another. Truly open models would just build on what is publicly available.

                                  • yesterday at 6:25 PM

                                    • alightsoul

                                      yesterday at 6:28 PM

                                      Crypto bros took the idea with some blockchain shit and no one takes it seriously anymore so it died

                                        • embedding-shape

                                          yesterday at 6:56 PM

                                          If the LLM/AI ecosystem starts actually needing some Person-To-Person (or maybe Agent-To-Agent?) payment system because things actually get smart enough to be useful autonomously, they're gonna need some way to send money/currency around. Depending on how banks will react to this need, we might see another return of digital currencies from the current winter.

                                            • idiotsecant

                                              today at 12:29 AM

                                              They used to say that cryptocurrency will be the dopamine layer of the first artificial intelligence

                                  • ignoramous

                                    yesterday at 6:32 PM

                                    UAE's IFM / LLM360 MO is indeed "fully open source" LLMs: https://www.llm360.ai/reports/LLM360-Towards-Fully-Transpare...

                                    • echelon

                                      yesterday at 5:05 PM

                                      Eventually we'll just construct 100% synthetic training data that can reliably reproduce pretrains and fine tunes.

                                      The first broadly useful fully open source models will do this.

                                      We already have open data / open code / open weights for some domain-specific cases, such as audio models trained on large open datasets, eg. Tacotron / LJSpeech from waaay back in the day, though that is certainly not SOTA anymore.

                                      Distillation could possibly be considered an early case of this as raw AI outputs are themselves not copyrightable unless humans enrich, filter, or transform them. Granted, that does not handle the cases where the outputs are sufficiently similar to copyrighted original works.

                                        • waffleiron

                                          yesterday at 6:00 PM

                                          Where does that synthetic data come from? Magically just started existing?

                                          • chaosharmonic

                                            yesterday at 6:08 PM

                                            But how much of that synthetic data still ultimately derives from non-open sources? You'd still have to ask what a clean room implementation ultimately is, depending on how granular or aggressive a large publisher wanted to get about it.

                                            That said, I don't necessarily disagree with you. Talkie[1] presents an interesting case for it being at least possible to do this entirely on public domain material.

                                            But even that used Claude somewhere in the course of its training pipeline (it's listed as a contributor on their GitHub), so again, how granular you want to get with that is still a question.

                                            [1] https://talkie-lm.com/chat

                                    • trvz

                                      yesterday at 4:44 PM

                                      Why? Sure, I’d prefer it, too, but this is just another GNU/Linux vs. macOS situation: most of us would prefer the first, but actually get shit done on the latter.

                                        • eikenberry

                                          yesterday at 8:16 PM

                                          Why do you make the worse choice and not use what you would prefer to use? You have been able to "get shit done" on Linux for nearly 30 years. Have the courage of your convictions.

                                          • didibus

                                            yesterday at 5:12 PM

                                            And that's why companies shouldn't fear opening up, but having both is still a net benefit.

                                            • zufallsheld

                                              yesterday at 4:59 PM

                                              Without open-source, there'd be no macOS.. So good thing, it exists.

                                              • homarp

                                                yesterday at 4:46 PM

                                                which is why everyone runs docker on mac, to get shit done.

                                                • verdverm

                                                  yesterday at 6:18 PM

                                                  we get shit done on the cloud with the former rather than the later

                                                  I personally find the analogy unconvincing, the UX dimension is completely different as I can use the same harness with any model; and the year of the linux desktop is coming soon (tm)

                                              • verdverm

                                                yesterday at 6:14 PM

                                                I believe Olmo from AllenAi is this

                                                https://allenai.org/olmo

                                                Open models can be used/changed for social manipulation too, by anyone, which scares a bunch of people, as opposed to the dark pattern manipulation from Big Ai/Tech

                                                • theplumber

                                                  yesterday at 10:50 PM

                                                  But that would be impossible due copyrights laws. If the law would apply Anthropic and OpenAI executives would be in jail

                                                  • cute_boi

                                                    yesterday at 5:44 PM

                                                    Money is the issue here, no one wants to fund it.

                                                      • __MatrixMan__

                                                        yesterday at 6:08 PM

                                                        I'm sure anthropic didn't want to fund the extra "safety" guardrails they put into fable, but they were forced to, else they couldn't release it.

                                                        Sure there are all kinds of problems with that situation. But it still demonstrates that they can be coerced: play nice or don't play at all.

                                                • a11r

                                                  yesterday at 5:05 PM

                                                  It is great to see another player introduce a fully open stack. Nvidia's Nemotron is the only other prominent one I know of.

                                                  All that said, the headline claims do not match the self-reported performance. For example, the dense 32B model is significantly behind Qwen3.8 27B (chart towards the bottom of https://ifm.ai/blog/k2). Gemma4 31B is not in the comparison set. This is the most important sweet spot for self hosted open-weight models today and real competition here will be very welcome.

                                                    • baron3dl

                                                      yesterday at 5:49 PM

                                                      https://allenai.org/ has the fully open olmo also

                                                      • xienze

                                                        yesterday at 5:13 PM

                                                        They have the 32B listed as "stage 1" with the note "final checkpoint to be released." So, not finished yet. Not sure why you'd release it if it's not finished, but that's the explanation.

                                                        The 7B does look very, very good however.

                                                          • bluejay2387

                                                            yesterday at 6:22 PM

                                                            In this case, the fully open source pipeline is probably as valuable or more so than the weights, so releasing early has some justification.

                                                            • WithinReason

                                                              yesterday at 5:34 PM

                                                              32B performs worse than the 7B model so I'm sure they will improve it

                                                      • cogman10

                                                        yesterday at 7:02 PM

                                                        My quick review of the 3.7B model (because I was interested) is that it's not to be trusted for coding.

                                                        It failed my basic test I like to ask models and generated incorrect code. When prompted about the bug, it preceded to start hallucinating non-existent APIs. After doing that it got caught in a loop trying to desk check the solution that didn't work.

                                                          • cogman10

                                                            yesterday at 10:01 PM

                                                            7B produced 2 answers, 1 was correct though more expensive and the second was incorrect.

                                                            The first attempt with 7B the model got stuck in an infinite loop.

                                                            • dotancohen

                                                              yesterday at 7:55 PM

                                                              I'd you have some tips for coming up with such tests, I would love to hear them. My Gmail username is the same as my HN username. Thank you!

                                                                • cogman10

                                                                  yesterday at 8:11 PM

                                                                  It's actually just a coding interview test that I liked to ask in the past. You can find it and others on leetcode.

                                                                  The reason I personally like my question is because it's pretty close to some of the real world work we do. It's mostly mundane and easy to bang out, but really easy for someone to do a n log n solution where an n solution exists.

                                                                  A good example (but not my question) would be something like

                                                                  "I have a list of People objects with a `first` and `last` name. Write a function which groups together all the People with the same last name in `your language of choice`"

                                                                    • dotancohen

                                                                      yesterday at 8:50 PM

                                                                      LLMs have a problem with that type of question? I might try it later at home.

                                                                        • cogman10

                                                                          yesterday at 8:54 PM

                                                                          Now a days? No. It's actually getting to be a bad question because they all push out about the exact same answer.

                                                                          But much earlier they did and, apparently, these really small models still do. At this point it serves as more of a smoke test for me. Success means little, failure means a lot.

                                                                            • dotancohen

                                                                              yesterday at 11:47 PM

                                                                              Terrific, thank you.

                                                              • xienze

                                                                yesterday at 8:28 PM

                                                                Not sure a model that small is really supposed to be used for any real coding. At that size you're usually using the model to do simple tasks like summarization.

                                                                  • cogman10

                                                                    yesterday at 8:33 PM

                                                                    To be clear, the question wasn't a complex one. It was more on the level of "could I use this for a fast inline coder" IE, single somewhat simple function question.

                                                                    I wouldn't have dreamed to use this as an agent model.

                                                                    7B models of the past have been able to pass this question. I've not tested it on a 4B model until now.

                                                            • RandyOrion

                                                              today at 3:48 AM

                                                              Good to see new open source/weight LLM families. Waiting for the final release of https://huggingface.co/IFM/K2-Horizon-32B .

                                                              • piinbinary

                                                                yesterday at 4:20 PM

                                                                A bit off topic, but I think I'm starting to get model fatigue. These come out 10x faster than new Javascript frameworks were coming out 10 years ago (at least new models are far easier to adopt).

                                                                  • hungryhobbit

                                                                    yesterday at 5:45 PM

                                                                    There was a time when every new PC CPU coming out was a giant deal: "Guys have you heard about this new Pentium processor, it's incredible?"

                                                                    But over time, more and more people got into the chip-making business, and the big players started releasing more and more chips. Now only the die-hard CPU trackers worry about every new CPU and exactly how it's better ... while everyone else just worries about "which CPU will be good enough at this moment".

                                                                    I think models are on that same arc.

                                                                      • pantelisk

                                                                        yesterday at 7:51 PM

                                                                        Same for smartphones. There was a time of unlimited hype and secrery around new iphones. Who remembers the story of an iphone 5 prototype left a bar. Journalists were going crazy, people were signing petitions for Apple to not hunt down but instead forgive the employee that made such a grave mistake. People were offering millions to buy the prototype so they can brag they got the new phone 2 weeks before everyone else did.

                                                                        Or who remembers the dancing disease of 1518, were people would stop what they are doing and start randomly doing the same dance. The lords? Out of their minds. The priests? Terrified the devil had taken hold of the flock! I have come to believe that it was probably some tik-tok like hype trend of doing a fortnite dance while waiting in line for bread and communion. And the energy back then, like now, was off the charts.

                                                                        Hype and memetic trend seeking encoded deep in human psyche.

                                                                    • wuhhh

                                                                      yesterday at 4:31 PM

                                                                      At least this one can claim being fully open to differentiate it

                                                                      • dgellow

                                                                        yesterday at 6:27 PM

                                                                        Honestly, you don’t have to pay attention. What you do with models matters way more than the models themselves, and you don’t need frontier for the vast, vast majority of use cases

                                                                        • kelseyfrog

                                                                          yesterday at 4:31 PM

                                                                          Just wait until RSI gains enough traction. We'll be compute-limited rather than labor-limited.

                                                                            • JSR_FDED

                                                                              yesterday at 6:55 PM

                                                                              Repetitive Strain Injury inverts this statement

                                                                              • yesterday at 5:17 PM

                                                                        • cesarvarela

                                                                          yesterday at 6:25 PM

                                                                          I find it funny that while these releases are a technological miracle, the charts in the doc use tiny fonts and are hard to read. Goes with the idea that coding might be solved, but taste isn't.

                                                                            • mzmzmzm

                                                                              yesterday at 8:13 PM

                                                                              Accessibility isn't "solved," but there are certainly standards for things like color contrast. Maybe inbetween taste and coding there are better targets still being missed.

                                                                                • culi

                                                                                  yesterday at 10:07 PM

                                                                                  In fact, automated a11y checks and tooling is quite advanced nowadays and tragically underutilized by web developers. Now that we have llms to scale all the shitty code of front-end devs at startups, I feel increasingly hopeless about things ever improving

                                                                          • uniclaude

                                                                            yesterday at 5:06 PM

                                                                            Seeing this the day all major closed LLMs went offline is quite the reminder of how valuable open source can be.

                                                                              • OmniCrativeWorx

                                                                                yesterday at 6:49 PM

                                                                                [flagged]

                                                                            • jon9544hn

                                                                              yesterday at 4:23 PM

                                                                              Here’s the link (K2)[https://ifm.ai/k2/] as the originally linked link is a login url.

                                                                              • justin_

                                                                                yesterday at 7:06 PM

                                                                                I'm glad to see some development in the space of "truly open" models that share training data and other recipes. As the costs for hardware fall over time (hopefully), we should see more possibility in fine-tuning and developing software to inspect the source training material.

                                                                                Some other open models I'm aware of:

                                                                                   - OLMo
                                                                                   - Apertus
                                                                                   - Soofi
                                                                                   - OpenEuroLLM
                                                                                   - llm-jp
                                                                                
                                                                                OLMo is perhaps the most famous, and their Dolma training corpus has been reused in other projects. It looks like the K2 training materials haven't been released yet, but I'm interested to see what they did for training "long-horizon agentic tasks". I'm aware of SWE-smith + SWE-gym but I'm guessing there's a lot more out there now.

                                                                                I'm no expert, which is part of why these projects excite me. I'm hoping they can be good projects to learn from as well.

                                                                                • mmastrac

                                                                                  yesterday at 4:55 PM

                                                                                  The comparisons with other models here are odd.. the other models change depending on the task. It would be far more useful to at least compare against the more recent open models (DS4Flash/GLM53Flash/Qwen38).

                                                                                    • cogman10

                                                                                      yesterday at 6:33 PM

                                                                                      They are trying to keep the models within the same quant class, which is tough to do since a lot of models aren't distilled to lower quants.

                                                                                      There is, for example, no Qwen3.8 7B.

                                                                                      It is odd to me, though, that they didn't run the same benchmark suite for the various quants.

                                                                                  • kzrdude

                                                                                    yesterday at 9:27 PM

                                                                                    The Uno "diffusion adaptor" will take a while for me to understand, but sounds very interesting.

                                                                                    https://huggingface.co/IFM/K2-Horizon-7B-Uno

                                                                                    • kamranjon

                                                                                      yesterday at 4:20 PM

                                                                                      it's funny that the tagline is Radically Open, but you're immediately hit with http login - maybe this was the wrong link?

                                                                                        • sottol

                                                                                          yesterday at 4:24 PM

                                                                                          It's not the blog post, but there's some info here:

                                                                                          https://ifm.ai/k2/

                                                                                          375 A23B, 36 A4B, 32B, 7B, 3.7B, 0.9B variants.

                                                                                          > 32B: Ranking among the top models in its class, 32B is our most powerful dense model, balancing capability, adaptability, and local deployability.

                                                                                          > 7B: The industry’s best-performing model under 10B combines strong software engineering and expert knowledge in a package small enough to run on a phone.

                                                                                          • gs17

                                                                                            yesterday at 4:21 PM

                                                                                            https://ifm.ai/k2/ seems to work for me.

                                                                                              • esafak

                                                                                                yesterday at 4:24 PM

                                                                                                But it's missing the all-important charts that the blog had before it started asking for authentication.

                                                                                                  • wmedrano

                                                                                                    yesterday at 4:25 PM

                                                                                                    You can find some of the charts on huggingface

                                                                                                    https://huggingface.co/collections/IFM/k2-horizon

                                                                                                      • sottol

                                                                                                        yesterday at 4:29 PM

                                                                                                        Thanks! Qwen-3.8 27B seems to benchmark better but I'd like to try this some time.

                                                                                                          • verdverm

                                                                                                            yesterday at 6:21 PM

                                                                                                            little qwen is my favorite for the homelab, vllm 0.28 now supports the dflash2 to go with it

                                                                                        • throwawayffffas

                                                                                          yesterday at 11:22 PM

                                                                                          While the open approach is commendable. The 32b and 36b models are inferior to qwen 3.8 27b, at least according to benchmark numbers. I would have liked to have seen both compared to 27b, not only the dense one. Also would have liked to see more coding benchmarks in the full table.

                                                                                          • afzalive

                                                                                            yesterday at 4:45 PM

                                                                                            Not to be confused with Kimi K2. Out of all the names they could've used, they picked one that would be confusing.

                                                                                              • bee_rider

                                                                                                yesterday at 5:06 PM

                                                                                                I kind of assumed all the K2 names were puns. K2 is quite tall, so to get to the top of it you have to be really good at hill climbing. Anyway it’s a pretty well known mountain so I don’t think anyone can call dibs on it.

                                                                                                • Topfi

                                                                                                  yesterday at 8:24 PM

                                                                                                  Not to be confused itself with K2 Think by MBZUAI...

                                                                                              • TechSquidTV

                                                                                                yesterday at 6:59 PM

                                                                                                I attempted their chat demo to see the speed and it stated the model couldnt be found.

                                                                                                edit: Tried signing up and using the internal playground. Holy shit thats fast.

                                                                                                • sottol

                                                                                                  yesterday at 4:25 PM

                                                                                                  The press release: https://ifm.ai/k2/press-release/

                                                                                                    • reasonableklout

                                                                                                      yesterday at 6:51 PM

                                                                                                      Interesting, have not heard of this company/org before. It seems they're from a UAE university?

                                                                                                  • luckydata

                                                                                                    yesterday at 4:28 PM

                                                                                                    both repositories for pre-training and post-training are actually empty... someone might have jumped the gun on the release.

                                                                                                      • verdverm

                                                                                                        yesterday at 6:24 PM

                                                                                                        [dead]

                                                                                                    • prometheus1992

                                                                                                      yesterday at 5:18 PM

                                                                                                      Nice! can't wait to add these in my local stack and try them out.

                                                                                                      • villish

                                                                                                        yesterday at 5:48 PM

                                                                                                        Frontier. Everything is frontier. K2 not to be confused with the other K2, or K3 that is also frontier.

                                                                                                        • yesterday at 5:59 PM

                                                                                                          • luciana1u

                                                                                                            yesterday at 6:07 PM

                                                                                                            i'll believe 'radically open' when the training data ships alongside the weights. until then it's a very fast demo.

                                                                                                              • adrian_b

                                                                                                                yesterday at 6:15 PM

                                                                                                                I just looked on Huggingface.co, and the training data is there.

                                                                                                                For example, 3.3 Tbyte for code reasoning, 4.5 Tbyte for mathematical reasoning, 8.4 Tbyte of pre-train behaviors, and so on.

                                                                                                                I did not compute the sum of the dataset sizes, but it appears to be some tens of Tbyte. Nonetheless, I assume that this amount of training data is more than an order of magnitude less than what OpenAI, Anthropic and the like have used, which must have been at least many hundreds of Tbyte, but more likely several thousands of Tbyte of data.

                                                                                                              • dakolli

                                                                                                                yesterday at 6:10 PM

                                                                                                                Hey its a lot mpre thsn Anthropic which you probably use everyday all day without complaints.

                                                                                                                  • luciana1u

                                                                                                                    yesterday at 10:28 PM

                                                                                                                    [dead]