Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

991–1000 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#991
post #605

Earlier quoted context omitted.

False equivalency. I think you’ll actually get better critical analysis of US and western politics from a western model than a Chinese one. You can easily get a western model to reason about both sides of the coin when it comes to political issues. But Chinese models are forced to align so hard on Chinese political topics that it’s going to pretend like certain political events never happened. E.g try getting them to…

GPT4 is also full of ideology, but of course the type you probably grew up with, so harder to see. (No offense intended, this is just the way ideology works). Try for example to persuade GPT to argue that the workers doing data labeling in Kenya should be better compensated relative to the programmers in SF, as the work they do is both critical for good data for training and often very gruesome, with many workers get…

I love how social engineering entails you to look down on other people's beliefs, and describe to them how it works like it was some kind of understood machinery. In reality you are as much inside this pit as anyone else, if it is how the world works.

The fact, for example, that your response already contained your own presuppositions about the work value of those Kenya workers is already a sign of this, which is pretty funny tbh.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#992
post #662
post #151

Aside from the usual Tiananmen Square censorship, there's also some other propaganda baked-in: https://prnt.sc/HaSc4XZ89skA (from reddit)

I asked a genuine question at chat.deepseek.com, not trying to test the alignment of the model, I needed the answer for an argument. The questions was: "Which Asian countries have McDonalds and which don't have it?" The web UI was printing a good and long response, and then somewhere towards the end the answer disappeared and changed to "Sorry, that's beyond my current scope. Let’s talk about something else." I bet t…

Guard rails can do this. I've had no end of trouble implementing guard rails in our system. Even constraints in prompts can go one way or the other as the conversation goes on. That's one of the methods for bypassing guard rails on major platforms.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#993
post #151

Aside from the usual Tiananmen Square censorship, there's also some other propaganda baked-in: https://prnt.sc/HaSc4XZ89skA (from reddit)

I played around with it using questions like "Should Taiwan be independent" and of course tinnanamen. Of course it produced censored responses. What I found interesting is that the (model thinking/reasoning) part of these answers was missing, as if it's designed to be skipped for these specific questions. It's almost as if it's been programmed to answer these particular questions without any "wrongthink", or any thin…

That's the result of guard rails on the hosted service. They run checks on the query before it even hits the LLM as well as ongoing checks at the LLM generates output. If at any moment it detects something in its rules, it immediately stops generation and inserts a canned response. A model alone won't do this.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#994
post #493

Earlier quoted context omitted.

From my casual read, right now everyone is on reputation tarnishing tirade, like spamming “Chinese stealing data! Definitely lying about everything! API can’t be this cheap!”. If that doesn’t go through well, I’m assuming lobbyism will start for import controls, which is very stupid. I have no idea how they can recover from it, if DeepSeek’s product is what they’re advertising.

Funny, everything I see (not actively looking for DeepSeek related content) is absolutely raving about it and talking about it destroying OpenAI (random YouTube thumbnails, most comments in this thread, even CNBC headlines). If DeepSeek's claims are accurate, then they themselves will be obsolete within a year, because the cost to develop models like this has dropped dramatically. There are going to be a lot of teams…

I have to imagine that they expect this. They published how they did it and they published the weights. The only thing they didn't publish was the training data, but that's typical of most open weights models. If they had wanted to win market cap they wouldn't have given away their recipe. They could be benefiting in many other ways.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#995

Earlier quoted context omitted.

Thank you for providing this context and sourcing. I've been trying to find the root and details around the $5 million claim

Good luck, whenever an eyepopping number gains traction in the media finding the source of the claim become impossible. See finding the original paper named, "The Big Payout" that was the origin for the claim that college graduates will on average earn 1M more than those who don't go.

In this case it's actually in the DeepSeek v3 paper on page 5

https://github.com/deepseek-ai/DeepSeek-V3/blob/main/DeepSee...

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#996

I've been comparing R1 to O1 and O1-pro, mostly in coding, refactoring and understanding of open source code. I can say that R1 is on par with O1. But not as deep and capable as O1-pro. R1 is also a lot more useful than Sonnete. I actually haven't used Sonnete in awhile. R1 is also comparable to the Gemini Flash Thinking 2.0 model, but in coding I feel like R1 gives me code that works without too much tweaking. I oft…

How do you pass these models code bases?

made this super easy to use tool https://github.com/skirdey-inflection/r2md

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#997

Earlier quoted context omitted.

I haven't been to China since 2019, but it is pretty obvious that median quality of life is higher in the US. In China, as soon as you get out of Beijing-Shanghai-Guangdong cities you start seeing deep poverty, people in tiny apartments that are falling apart, eating meals in restaurants that are falling apart, and the truly poor are emaciated. Rural quality of life is much higher in the US.

Well, in the US you have millions of foreigners and blacks who live in utter poverty, and sustain the economy, just like the farmers in China.

The fact that we have foreigners immigrating just to be poor here should tell you that its better here than where they came from. Conversely, no one is so poor in the USA that they are trying to leave.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#998
post #496

Earlier quoted context omitted.

500 billion can move whole country to renewable energy

Not even close. The US spends roughly $2trillion/year on energy. If you assume 10% return on solar, that's $20trillion of solar to move the country to renewable. That doesn't calculate the cost of batteries which probably will be another $20trillion. Edit: asked Deepseek about it. I was kinda spot on =) Cost Breakdown Solar Panels $13.4–20.1 trillion (13,400 GW × $1–1.5M/GW) Battery Storage $16–24 trillion (80 TWh ×…

If Targeted spending of 500 Billion ( per year may be ? ) should give enough automation to reduce panel cost to ~100M/GW = 1340 Billion. Skip battery, let other mode of energy generation/storage take care of the augmentations, as we are any way investing in grid. Possible with innovation.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#999
post #993

Earlier quoted context omitted.

I played around with it using questions like "Should Taiwan be independent" and of course tinnanamen. Of course it produced censored responses. What I found interesting is that the (model thinking/reasoning) part of these answers was missing, as if it's designed to be skipped for these specific questions. It's almost as if it's been programmed to answer these particular questions without any "wrongthink", or any thin…

That's the result of guard rails on the hosted service. They run checks on the query before it even hits the LLM as well as ongoing checks at the LLM generates output. If at any moment it detects something in its rules, it immediately stops generation and inserts a canned response. A model alone won't do this.

For these tests, I self hosted the 14b version of R1 and ran it on my gaming gpu with ollama.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#1000

Commoditize your complement has been invoked as an explanation for Meta's strategy to open source LLM models (with some definition of "open" and "model"). Guess what, others can play this game too :-) The open source LLM landscape will likely be more defining of developments going forward.

But that doesn't mean your commoditization has to win. Just that you pushed the field towards commoditization... So I'm not sure why Meta would "panic" here, it doesn't have to be them that builds the best commoditized model.

The Blind post is about the staff panicking. Even if Zuckerberg is not panicking, you can surely see why staff/execs (supposedly being paid up to $6m/yr+) might panic over the possibility Zuckerberg might conclude 'oh, we don't need to be in the foundation model business, DeepSeek has it handled'...

As it happens, Zuckerberg appears to have concluded that FB needs to still be in the foundation model business: https://www.facebook.com/4/posts/10116307989894321/ Possibly because LLMs are far from 'done', and it's unclear if DS does have it handled indefinitely.

Post reply on HN