Live data from Hacker News

K2 Horizon: A connected fleet of six open models

ifm.ai

131–140 of 144 posts

Re: K2 Horizon: A connected fleet of six open models

#131
post #5

Fully open models really need to be a big part of the AI future. That includes all source code, open training data, how it's organized, fed to the model, processed, etc. Until that becomes a thing you're always going to be left wondering what exactly lies underneath the closed model you are using, leaving open the possibility for societal manipulation.

I believe Olmo from AllenAi is this https://allenai.org/olmo Open models can be used/changed for social manipulation too, by anyone, which scares a bunch of people, as opposed to the dark pattern manipulation from Big Ai/Tech

No, they just tell you what data they used, it still includes e.g. Common Crawl. Not Open, just willing to state what they fed into the training.

Re: K2 Horizon: A connected fleet of six open models

#133
post #81

Earlier quoted context omitted.

Other than open training data (currently legally impossible), all of this holds for basically every major Chinese-made model. They not only open the weights but publish detailed methodology papers alongside the models in arXiv and even open source the code.

Inference code, yes, but the specifics of their training process (as well as the training of the vast majority of all other open weights models) are still a complete blackbox, and I can't think of any Chinese model that made its training corpus public.

This is absolutely not true. DeepSeek is most famous for publishing really in-depth papers on their training process but the other labs have started to do the same as well.

If by "training corpus" you mean the actual data I already acknowledged that that's currently legally impossible.

In fact everything I just said I said in my original comment. It's like you didn't read it at all.

DeepSeek's GRPO Infrastructure, multi-stage training pipeline, and their "cold start" phase have been massively influential in LLM research.

Re: K2 Horizon: A connected fleet of six open models

#134
post #81

Earlier quoted context omitted.

Other than open training data (currently legally impossible), all of this holds for basically every major Chinese-made model. They not only open the weights but publish detailed methodology papers alongside the models in arXiv and even open source the code.

They don’t release all the code.

DeepSeek has released a ton of low-level AI infrastructure code, libraries, and mathematical models on GitHub. Their tools and agent environments are fully open sourced including their harness. They release complete PyTorch and Hugging-face compatible python files detailing their configuration, tokenizers, and layers of the architecture.

The only thing they don't release is their data-filtering pipelines but they detail even that in their public-access papers.

DeepSeek is truly as open source as you can possibly legally get. Besides the data itself, it's completely reproducible by anyone else.

I don't think americans yet acknowledge just how radically transparent Chinese labs are being (and how much even the west benefits from it).

Re: K2 Horizon: A connected fleet of six open models

#135
post #133

Earlier quoted context omitted.

Inference code, yes, but the specifics of their training process (as well as the training of the vast majority of all other open weights models) are still a complete blackbox, and I can't think of any Chinese model that made its training corpus public.

This is absolutely not true. DeepSeek is most famous for publishing really in-depth papers on their training process but the other labs have started to do the same as well. If by "training corpus" you mean the actual data I already acknowledged that that's currently legally impossible. In fact everything I just said I said in my original comment. It's like you didn't read it at all. DeepSeek's GRPO Infrastructure, mu…

> If by "training corpus" you mean the actual data I already acknowledged that that's currently legally impossible.

Did you read the site this very post links to? The entire point is that the training corpus, recipe, and scripts, as well as intermediate checkpoints, will be made available for K2 Horizon. Chinese models these days don't even release pre-trained weights anymore; all you get now is the finished post-trained product.

I've read your comment. I'm doubting you've even read the thing you were commenting on.

Re: K2 Horizon: A connected fleet of six open models

#136
post #121
post #111

Earlier quoted context omitted.

> for their heirs and dependents Why can't they do what the rest of us do? Earn and save money during your working life and leave _that_ for your heirs. Let copyright die with the author.

Work is often only recently published when an artist or author dies but has taken years of non-earning to create. I know this them-and-us thinking is fashionable in the tech world but the reality is that the majority of creative people don’t earn much and never have, and copyright was developed not to give them extra power over the rest of us but to create a framework for creative work to earn them an income at all.…

Sure, some number of years is reasonable - but what do you think that is? Because 70 is insane. A single bestseller should not be able to support an extended family over three generations; at some point the rent-seeking becomes excessive. If you publish a work, at some point it stops being yours. The fact that this takes a whole lifetime is already very generous.

Re: K2 Horizon: A connected fleet of six open models

#137
post #136
post #121

Earlier quoted context omitted.

Work is often only recently published when an artist or author dies but has taken years of non-earning to create. I know this them-and-us thinking is fashionable in the tech world but the reality is that the majority of creative people don’t earn much and never have, and copyright was developed not to give them extra power over the rest of us but to create a framework for creative work to earn them an income at all.…

Sure, some number of years is reasonable - but what do you think that is? Because 70 is insane. A single bestseller should not be able to support an extended family over three generations; at some point the rent-seeking becomes excessive. If you publish a work, at some point it stops being yours. The fact that this takes a whole lifetime is already very generous.

I don't disagree; personally I think you could go with life plus 35 and cap the whole thing at 80 years from publication.

But I do think some potential post-mortem protection is essential for creative work to remain viable, and that means that any post-mortem buyer of an artist's estate has to be able to get value from recent work for a period of time.

This whole discussion is somewhat fantastical now anyway, because copyright is fucked.

But the intent was always to make working artists' lives possible; the various copyright extensions have always been for the benefit of corporate copyright holders, and it is unfair to vilify individual working artists for that.

Re: K2 Horizon: A connected fleet of six open models

#138
post #117

Earlier quoted context omitted.

That sounds dystopian to me and is against the hacker ethics.

Letting information that want to be free, be free, sounds exactly like the hacker ethics to myself, and the link you shared earlier would agree.

The last point is pretty clear, and the text below clarifies it:

> To protect the privacy of the individual and to strengthen the freedom of the information which concern the public the yet last point was added.

The privacy of individuals is important, regardless where they store their private data. Their account information -- what they buy, their medical information and so on is stored on servers and could be hacked.

Re: K2 Horizon: A connected fleet of six open models

#140
post #138

Earlier quoted context omitted.

Letting information that want to be free, be free, sounds exactly like the hacker ethics to myself, and the link you shared earlier would agree.

The last point is pretty clear, and the text below clarifies it: > To protect the privacy of the individual and to strengthen the freedom of the information which concern the public the yet last point was added. The privacy of individuals is important, regardless where they store their private data. Their account information -- what they buy, their medical information and so on is stored on servers and could be hacke…

> The last point is pretty clear

I think the second point is equally clear, and further up on the list.

I agree that what people buy, their medical information and so on should be private, hence it should only be offline and not stored/handled on computers connected to the internet at all, the internet should be for public data exclusively, is my argument in the initial comment. Medical information would be only on effectively airgapped computers, as that data should be private, as you say.

Post reply on HN