Live data from Hacker News

Deepseek R1-0528

huggingface.co

151–160 of 264 posts

Re: Deepseek R1-0528

#151

Earlier quoted context omitted.

It's. not. open. source! https://www.downloadableisnotopensource.org/

Open source is a crazy new beast in the AI/ML world. We have numerous artifacts to reason about: - The model code - The training code - The fine tuning code - The inference code - The raw training data - The processed training data (which might vary across various stages of pre-training and potentially fine-tuning!) - The resultant weights - The inference outputs (which also need a license) - The research papers (hop…

I'd argue we don't need a 10 star system. The single bit we have now is enough. And the question is also pretty clear: did $company steal other peoples work?

The answer is also known. So the reason one would want an open source model (read reproducible model), would be that of ethics

Re: Deepseek R1-0528

#152
post #101

Earlier quoted context omitted.

I don't think people make the distinction like that. The open source vs non open source distinction boils down to, usually, can you use it for commercial use. what you're saying is just that it's non reproducible, which is a completely valid but separate issue

But where's the source? I just see a binary blob, what makes it open source?

You can fine-tune their weights and release your own take.

E.g. see all the specialized third-party models out there based on Qwen.

"Open-source" is the wrong word here, what they mean is "you can modify and redistribute these weights".

Re: Deepseek R1-0528

#153
post #4

I love how Deepseek just casually drops new updates (that deliver big improvements) without fanfare.

Much more preferred to what OpenAI always did and Anthropic recently started doing. Just write some complicated narrative about how scary this new model is and how it tried to escape and deceive and hack the mainframe while telling the alignment operators bed time stories.

Really? I missed this. The new hype trick is implying the new LLM releases are almost AGI? Love it.

Re: Deepseek R1-0528

#154

Earlier quoted context omitted.

No sign of what source material it was trained on though right? So open weight rather than reproducible from source. I remember there's a project "Open R1" that last I checked was working on gathering their own list of training material, looks active but not sure how far along they've gotten: https://github.com/huggingface/open-r1

> No sign of what source material it was trained on though right? out of curiosity, does anyone do anything "useful" with that knowledge? it's not like people can just randomly train models..

Many are speculating it was trained by o1/o3 for some of the initial reasoning.

Re: Deepseek R1-0528

#155

Earlier quoted context omitted.

No sign of what source material it was trained on though right? So open weight rather than reproducible from source. I remember there's a project "Open R1" that last I checked was working on gathering their own list of training material, looks active but not sure how far along they've gotten: https://github.com/huggingface/open-r1

> No sign of what source material it was trained on though right? out of curiosity, does anyone do anything "useful" with that knowledge? it's not like people can just randomly train models..

Are there any widely used models that publish this? If not, then no I guess.

Re: Deepseek R1-0528

#156
post #151

Earlier quoted context omitted.

Open source is a crazy new beast in the AI/ML world. We have numerous artifacts to reason about: - The model code - The training code - The fine tuning code - The inference code - The raw training data - The processed training data (which might vary across various stages of pre-training and potentially fine-tuning!) - The resultant weights - The inference outputs (which also need a license) - The research papers (hop…

I'd argue we don't need a 10 star system. The single bit we have now is enough. And the question is also pretty clear: did $company steal other peoples work? The answer is also known. So the reason one would want an open source model (read reproducible model), would be that of ethics

I truthfully cannot think of a single model that satisfies your criteria.

And if we wait for the the internet to be wholly eaten by AI, if we accept perfect as the enemy of good, then we'll have nothing left to cling to.

> And the question is also pretty clear: did $company steal other peoples work?

Who the hell cares? By the time this is settled - and I'd argue you won't get a definitive agreement - the internet will be won by the hyperscalers.

Accept corporate gifts of AI, and keep pushing them forward. Commoditize. Let there be no moat.

There will be infinite synthetic data available to us in the future anyway. And none of this bickering will have even mattered.

Re: Deepseek R1-0528

#158

Earlier quoted context omitted.

OpenAI does a lot of work hyping themselves up and creating buzz around things they do or have a vague idea that they might try to do in the future. Not to make people aware of GenAI, but to make sure OpenAI continues to be perceived as the AI company. The company that leads and revolutionizes, with everyone just copying them and trying to match them. That perception is a significant part of their value and probably…

Considering just how quickly others followed it's also obviously not the case. Infact the best AI software as in most useful is not theirs. Claude is far more reliable.

Only goes to shown how strong their brand already is in the eyes of the public. I question how much active maintenance it takes, though. In my eyes, they already won big with ChatGPT - the name still is synonymous with LLMs to the general population. Everyone knows what ChatGPT is. Few known what Claude or Gemini is; arguably, more people know what Deepseek is thanks to the splash they made tanking NVidia stock and becoming part of general news coverage for a few days. Still, for regular folks (including business folks in tech industry, too), they all are, respectively, "that other ChatGPT", "ChatGPT from Google" and "that Chinese ChatGPT". It's a pretty sticky perception.

Re: Deepseek R1-0528

#159
post #151

Earlier quoted context omitted.

Open source is a crazy new beast in the AI/ML world. We have numerous artifacts to reason about: - The model code - The training code - The fine tuning code - The inference code - The raw training data - The processed training data (which might vary across various stages of pre-training and potentially fine-tuning!) - The resultant weights - The inference outputs (which also need a license) - The research papers (hop…

I'd argue we don't need a 10 star system. The single bit we have now is enough. And the question is also pretty clear: did $company steal other peoples work? The answer is also known. So the reason one would want an open source model (read reproducible model), would be that of ethics

We use pop-cultural references to communicate all the time these days. Those don't necessarily come from only the most commonly known sections of these works, so the AI would necessarily need the full work (or a functional transformation of the work) to be able to hit the theoretical maximum of the ability to decode about and reason using such references. To exclude copyrighted works from the training set is to expect it to decode from the outside what amounts to humanity's own in-group jokes.

That's my formal argument. The less formal one is that copyright protection is something that smaller artists deserve more than rich conglomerates, and even then, durations shouldn't be "eternity and a day". A huge chunk of what is being "stolen" should be in the commons anyway.

Re: Deepseek R1-0528

#160
post #143

Earlier quoted context omitted.

> No sign of what source material it was trained on though right? out of curiosity, does anyone do anything "useful" with that knowledge? it's not like people can just randomly train models..

When you're trully open source, you can make ethings like this: Today we introduce OLMoTrace , a one-of-a-kind feature in the Ai2 Playground that lets you trace the outputs of language models back to their full, multi-trillion-token training data in real time. OLMoTrace is a manifestation of Ai2’s commitment to an open ecosystem – open models, open data, and beyond. https://allenai.org/blog/olmotrace

you can do these same, except you would need to be a pirate website. It would even be better. except illegal. but it would be better.
Post reply on HN