Live data from Hacker News

K2 Horizon: A connected fleet of six open models

ifm.ai

41–50 of 144 posts

Re: K2 Horizon: A connected fleet of six open models

#42
post #5

Fully open models really need to be a big part of the AI future. That includes all source code, open training data, how it's organized, fed to the model, processed, etc. Until that becomes a thing you're always going to be left wondering what exactly lies underneath the closed model you are using, leaving open the possibility for societal manipulation.

I believe Olmo from AllenAi is this

https://allenai.org/olmo

Open models can be used/changed for social manipulation too, by anyone, which scares a bunch of people, as opposed to the dark pattern manipulation from Big Ai/Tech

Re: K2 Horizon: A connected fleet of six open models

#43

[flagged]

I just looked on Huggingface.co, and the training data is there.

For example, 3.3 Tbyte for code reasoning, 4.5 Tbyte for mathematical reasoning, 8.4 Tbyte of pre-train behaviors, and so on.

I did not compute the sum of the dataset sizes, but it appears to be some tens of Tbyte. Nonetheless, I assume that this amount of training data is more than an order of magnitude less than what OpenAI, Anthropic and the like have used, which must have been at least many hundreds of Tbyte, but more likely several thousands of Tbyte of data.

Re: K2 Horizon: A connected fleet of six open models

#44
post #15
post #5

Fully open models really need to be a big part of the AI future. That includes all source code, open training data, how it's organized, fed to the model, processed, etc. Until that becomes a thing you're always going to be left wondering what exactly lies underneath the closed model you are using, leaving open the possibility for societal manipulation.

Why? Sure, I’d prefer it, too, but this is just another GNU/Linux vs. macOS situation: most of us would prefer the first, but actually get shit done on the latter.

we get shit done on the cloud with the former rather than the later

I personally find the analogy unconvincing, the UX dimension is completely different as I can use the same harness with any model; and the year of the linux desktop is coming soon (tm)

Re: K2 Horizon: A connected fleet of six open models

#45
post #12

Earlier quoted context omitted.

You can find some of the charts on huggingface https://huggingface.co/collections/IFM/k2-horizon

Thanks! Qwen-3.8 27B seems to benchmark better but I'd like to try this some time.

little qwen is my favorite for the homelab, vllm 0.28 now supports the dflash2 to go with it

Re: K2 Horizon: A connected fleet of six open models

#46
post #27
post #23

It is great to see another player introduce a fully open stack. Nvidia's Nemotron is the only other prominent one I know of. All that said, the headline claims do not match the self-reported performance. For example, the dense 32B model is significantly behind Qwen3.8 27B (chart towards the bottom of https://ifm.ai/blog/k2 ). Gemma4 31B is not in the comparison set. This is the most important sweet spot for self host…

They have the 32B listed as "stage 1" with the note "final checkpoint to be released." So, not finished yet. Not sure why you'd release it if it's not finished, but that's the explanation. The 7B does look very, very good however.

In this case, the fully open source pipeline is probably as valuable or more so than the weights, so releasing early has some justification.

Re: K2 Horizon: A connected fleet of six open models

#49
post #40
post #18

Earlier quoted context omitted.

The training data would need to have a permissive license for this to be possible.

Hear me out. Decentralized unstoppable storage, combined with decentralized unstoppable training, sorta like SETI for AI training. The seed of this tech already exists with IPFS and others like it. We know (some? all?) of the big labs have skirted copyright laws at one point or another. Truly open models would just build on what is publicly available.

[deleted]

Re: K2 Horizon: A connected fleet of six open models

#50
post #18
post #5

Fully open models really need to be a big part of the AI future. That includes all source code, open training data, how it's organized, fed to the model, processed, etc. Until that becomes a thing you're always going to be left wondering what exactly lies underneath the closed model you are using, leaving open the possibility for societal manipulation.

The training data would need to have a permissive license for this to be possible.

Or, we just need to get this over with and declare any digital data findable via the internet to just be public property of everyone. Everything becomes public, besides stuff you keep locally, and there is no difference anymore, it's all just data anyone can use for whatever. A 1 year grace period for everyone to pull stuff off they don't want to be a part of this bright new open era, then we just scrap everything related to intellectual property, copyright and similar stupid stuff, and slap UBI on top of all of it for good measure.
Post reply on HN