Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

921–930 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#921
post #662
post #151

Aside from the usual Tiananmen Square censorship, there's also some other propaganda baked-in: https://prnt.sc/HaSc4XZ89skA (from reddit)

I asked a genuine question at chat.deepseek.com, not trying to test the alignment of the model, I needed the answer for an argument. The questions was: "Which Asian countries have McDonalds and which don't have it?" The web UI was printing a good and long response, and then somewhere towards the end the answer disappeared and changed to "Sorry, that's beyond my current scope. Let’s talk about something else." I bet t…

Try again may be, it had no problem answering this for me.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#922
post #84

Interacting with this model is just supplying your data over to an adversary with unknown intents. Using an open source model is subjecting your thought process to be programmed with carefully curated data and a systems prompt of unknown direction and intent.

Open source means you set the system prompt.

But not the training data

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#923
post #119

Earlier quoted context omitted.

Seeing what china is doing to the car market, I give it 5 years for China to do to the AI/GPU market to do the same. This will be good. Nvidia/OpenAI monopoly is bad for everyone. More competition will be welcome.

That is not going to happen without currently embargo'ed litography tech. They'd be already making more powerful GPUs if they could right now.

Chinese companies are working euv litho, its coming.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#924
post #650

Earlier quoted context omitted.

Because literally before o1, no one is doing COT style test time scaling. It is a new paradigm. The talking point back then, is the LLM hits the wall. R1's biggest contribution IMO, is R1-Zero, I am fully sold with this they don't need o1's output to be as good. But yeah, o1 is still the herald.

I don't think Chain of Thought in itself was a particularly big deal, honestly. It always seemed like the most obvious way to make AI "work". Just give it some time to think to itself, and then summarize and conclude based on its own responses. Like, this idea always seemed completely obvious to me, and I figured the only reason why it hadn't been done yet is just because (at the time) models weren't good enough. (So…

But the longer you allocate tokens to CoT, the better it at solving the problem is a revolutionary idea. And model self correct within its own CoT is first brought out by o1 model.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#925
post #654

Earlier quoted context omitted.

Because literally before o1, no one is doing COT style test time scaling. It is a new paradigm. The talking point back then, is the LLM hits the wall. R1's biggest contribution IMO, is R1-Zero, I am fully sold with this they don't need o1's output to be as good. But yeah, o1 is still the herald.

Chain of Thought was known since 2022 ( https://arxiv.org/abs/2201.11903 ), we just were stuck in a world where we were dumping more data and compute at the training instead of looking at other improvements.

CoT is a common technique, but scaling law of more test time compute on CoT generation, correlates with problem solving performance is from o1.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#926
post #605

Earlier quoted context omitted.

False equivalency. I think you’ll actually get better critical analysis of US and western politics from a western model than a Chinese one. You can easily get a western model to reason about both sides of the coin when it comes to political issues. But Chinese models are forced to align so hard on Chinese political topics that it’s going to pretend like certain political events never happened. E.g try getting them to…

GPT4 is also full of ideology, but of course the type you probably grew up with, so harder to see. (No offense intended, this is just the way ideology works). Try for example to persuade GPT to argue that the workers doing data labeling in Kenya should be better compensated relative to the programmers in SF, as the work they do is both critical for good data for training and often very gruesome, with many workers get…

The Kenyan government isn't particularly in favor of this, because they don't want their essential workers (like doctors and civil servants) all quitting to become high-paid data labellers.

Unfortunately, one kind of industrial policy you might want to do attract foreign investment (like building factories) is to prevent local wages from growing too fast.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#927

Everyone is trying to say its better than the biggest closed models. It feels like it has parity, but its not the clear winner. But, its free and open and the quant models are insane. My anecdotal test is running models on a 2012 mac book pro using CPU inference and a tiny amount of RAM. The 1.5B model is still snappy, and answered the strawberry question on the first try with some minor prompt engineering (telling i…

* Yes I am aware I am not running R1, and I am running a distilled version of it.

If you have experience with tiny ~1B param models, its still head and shoulders above anything that has come before. IMO there have not been any other quantized/distilled/etc models as good at this size. It would not exist without the original R1 model work.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#928
I wonder if a language model can be treated as a policy over token-level actions instead of full response actions. Then each response from the language model is a full rollout of the policy. In math and coding, the reward for the response can be evaluated. This is not how DeepSeek works now, right? It treats full responses from the language model as the action if I understand correctly.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#929
post #586

Earlier quoted context omitted.

Humanity keeps running into copyright issues with every major leap in IT technology (photocopiers, tape cassettes, personal computers, internet, and now AI). I think it's about time for humanity to rethink their take on the unnatural restriction of information. I personally hope that countries recognize copyright and patents for what they really are and abolish them. Countries that refuse to do so can play catch up.

This is based on a flawed view of how we humans behave. Without incentive no effort. This is also the reason why socialism has and always will fail. People who put massive effort in creating original content need to be able to earn the rewards.

The premise, that forgoing copyright would necessitate the forgoing of incentives and rewards, is one entirely of your own assertion and was not implied in my above comment. I agree that your assertion is flawed.

There can be, and are, incentives and rewards associated with sharing information without flawed artificial constraints like copyright.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#930

For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…

Except it refuses to talk about the 1989 Tiananmen Square protests and massacre[0]. Are we really praising a model that is so blatantly censored by an authoritarian government?

[0]https://en.wikipedia.org/wiki/1989_Tiananmen_Square_protests...

Post reply on HN