Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

281–290 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#281
post #211

Earlier quoted context omitted.

I would think Meta - who open source their model - would be less freaked out than those others that do not.

The criticism seems to mostly be that Meta maintains very expensive cost structure and fat organisation in the AI. While Meta can afford to do this, if smaller orgs can produce better results it means Meta is paying a lot for nothing. Meta shareholders now need to ask the question how many non-productive people Meta is employing and is Zuck in the control of the cost.

That makes sense. I never could see the real benefit for Meta to pay a lot to produce these open source models (I know the typical arguments - attracting talent, goodwill, etc). I wonder how much is simply LeCun is interested in advancing the science and convinced Zuck this is good for company.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#282
post #151

Aside from the usual Tiananmen Square censorship, there's also some other propaganda baked-in: https://prnt.sc/HaSc4XZ89skA (from reddit)

Who cares? I ask O1 how to download a YouTube music playlist as a premium subscriber, and it tells me it can't help. Deepseek has no problem.

Oh wow, o1 really refuses to answer that, even though the answer that Deepseek gives is really tame (and legal in my jurisdiction): use software to record what's currently playing on your computer, then play stuff in the YTM app.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#283
post #182

Earlier quoted context omitted.

Correct me if I'm wrong but if Chinese can produce the same quality at %99 discount, then the supposed $500B investment is actually worth $5B. Isn't that the kind wrong investment that can break nations? Edit: Just to clarify, I don't imply that this is public money to be spent. It will commission $500B worth of human and material resources for 5 years that can be much more productive if used for something else - i.e…

The $500B is just an aspirational figure they hope to spend on data centers to run AI models, such as GPT-o1 and its successors, that have already been developed. If you want to compare the DeepSeek-R development costs to anything, you should be comparing it to what it cost OpenAI to develop GPT-o1 (not what they plan to spend to run it), but both numbers are somewhat irrelevant since they both build upon prior resea…

Thinking of the $500B as only an aspirational number is wrong. It’s true that the specific Stargate investment isn’t fully invested yet, but that’s hardly the only money being spent on AI development.

The existing hyperscalers have already sunk ungodly amounts of money into literally hundreds of new data centers, millions of GPUs to fill them, chip manufacturing facilities, and even power plants with the impression that, due to the amount of compute required to train and run these models, there would be demand for these things that would pay for that investment. Literally hundreds of billions of dollars spent already on hardware that’s already half (or fully) built, and isn’t easily repurposed.

If all of the expected demand on that stuff completely falls through because it turns out the same model training can be done on a fraction of the compute power, we could be looking at a massive bubble pop.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#284

Earlier quoted context omitted.

From my casual read, right now everyone is on reputation tarnishing tirade, like spamming “Chinese stealing data! Definitely lying about everything! API can’t be this cheap!”. If that doesn’t go through well, I’m assuming lobbyism will start for import controls, which is very stupid. I have no idea how they can recover from it, if DeepSeek’s product is what they’re advertising.

So you're saying that this is the end of OpenAI? Somehow I doubt it.

Hah I agree, they will find a way. In the end, the big winners will be the ones who find use cases other than a general chatbot. Or AGI, I guess.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#285

Earlier quoted context omitted.

> I guess censorship doesnt have as bad a rep in china as it has here It's probably disliked, just people know not to talk about it so blatantly due to chilling effects from aforementioned censorship. disclaimer: ignorant American, no clue what i'm talking about.

on the topic of censorship, US LLMs' censorship is called alignment. llama or ChatGPT's refusal on how to make meth or nuclear bombs is the same as not answering questions abput Tiananmen tank man as far as the matrix math word prediction box is concerned.

The distinction is that one form of censorship is clearly done for public relations purposes from profit minded individuals while the other is a top down mandate to effectively rewrite history from the government.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#286

Earlier quoted context omitted.

And (some people here are saying that)* if they are up-to-date is because they're cheating. The copium itt is astounding.

What’s the difference between what they do and what other ai firms do to openai in the us? What is cheating in a business context?

Chinese companies smuggling embargo'ed/controlled GPUs and using OpenAI outputs violating their ToS is considered cheating. As I see it, this criticism comes from a fear of USA losing its first mover advantage as a nation.

PS: I'm not criticizing them for it nor do I really care if they cheat as long as prices go down. I'm just observing and pointing out what other posters are saying. For me if China cheating means the GenAI bubble pops, I'm all for it. Plus no actor is really clean in this game, starting with OAI practically stealing all human content without asking for building their models.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#287

Earlier quoted context omitted.

Not necessarily if you are pushing against a data wall. One could ask: after adjusting for DS efficiency gains how much more compute has OpenAI spent? Is their model correspondingly better? Or even DS could easily afford more than $6 million in compute but why didn't they just push the scaling?

right except that r1 is demoing the path of approach for moving beyond the data wall

Can you clarify? How are they able to move beyond the data wall?

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#288
post #165

Earlier quoted context omitted.

In Communist theoretical texts the term "propaganda" is not negative and Communists are encouraged to produce propaganda to keep up morale in their own ranks and to produce propaganda that demoralize opponents. The recent wave of the average Chinese has a better quality of life than the average Westerner propaganda is an obvious example of propaganda aimed at opponents.

Is it propaganda if it's true?

Technically, as long as the aim/intent is to influence public opinion, yes. And most often it is less about being "true" or "false" and more about presenting certain topics in a one-sided manner or without revealing certain information that does not support what one tries to influence about. If you know any western media that does not do this, I would be very up to check and follow them, even become paid subscriber.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#289
post #268

I've been using https://chat.deepseek.com/ over My ChatGPT Pro subscription because being able to read the thinking in the way they present it is just much much easier to "debug" - also I can see when it's bending it's reply to something, often softening it or pandering to me - I can just say "I saw in your thinking you should give this type of reply, don't do that". If it stays free and gets better that's going to b…

The chain of thought is super useful in so many ways, helping me: (1) learn, way beyond the final answer itself, (2) refine my prompt, whether factually or stylistically, (3) understand or determine my confidence in the answer.
Post reply on HN