Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

971–980 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#971
post #647

Earlier quoted context omitted.

I mean US models are highly censored too.

How exactly? Is there any models that refuse to give answers about “the trail of tears”? False equivalency if you ask me. There may be some alignment to make the models polite and avoid outright racist replies and such. But political censorship? Please elaborate

I guess it depends on what you care about more: systemic "political" bias or omitting some specific historical facts.

IMO the first is more nefarious, and it's deeply embedded into western models. Ask how COVID originated, or about gender, race, women's pay, etc. They basically are modern liberal thinking machines.

Now the funny thing is you can tell DeepSeek is trained on western models, it will even recommend puberty blockers at age 10. Something I'm positive the Chinese government is against. But we're discussing theoretical long-term censorship, not the exact current state due to specific and temporary ways they are being built now.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#972

The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.

CEO of Scale said Deepseek is lying and actually has a 50k GPU cluster. He said they lied in the paper because technically they aren't supposed to have them due to export laws. I feel like this is very likely. They obvious did some great breakthroughs, but I doubt they were able to train on so much less hardware.

Only the DeepSeek V3 paper mentions compute infrastructure, the R1 paper omits this information, so no one actually knows. Have people not actually read the R1 paper?

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#973
post #566
post #480

Earlier quoted context omitted.

Reasoning from science fiction isn't a particularly strong approach. And every possible future is distopian - even the present is distopian in a practical sense. We have billions of people who live well below any standard I woudl consider acceptable.

Reasoning from science fiction is just stupid. A story first and foremost has to have conflict: if it doesn't there is no story, and thus all the stories have one. Science fiction also follows the anxieties of the time it is written in, as well as the conventions of the subgenre it's representing: i.e Star Trek doesn't have drones or remote surveillance really. Though it does accidentally have LLMs (via the concept o…

Great science fiction is grounded in conflict, as is human nature. There is a whole subtext of conflict in this, and other threads about AI: a future of machine oligarchs, of haves and have-nots. Great science fiction, like any great literature, is grounded in a deep understanding and a profound abstraction of humanity. I completely disagree that reasoning by science fiction is stupid, and the proof is in the pudding: science fiction writers have made a few great predictions.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#974

Earlier quoted context omitted.

I asked Chatgpt: how many civilians Israel killed in Gaza. Please provide a rough estimate. As of January 2025, the conflict between Israel and Hamas has resulted in significant civilian casualties in the Gaza Strip. According to reports from the United Nations Office for the Coordination of Humanitarian Affairs (OCHA), approximately 7,000 Palestinian civilians have been killed since the escalation began in October 2…

This accusation that American models are somehow equivalent in censorship to models that are subject to explicit government driven censorship is obviously nonsense, but is a common line parroted by astroturfing accounts looking to boost China or DeepSeek. Some other comment had pointed out that a bunch of relatively new accounts participating in DeepSeek related discussions here, on Reddit, and elsewhere are doing th…

is it really primarily an astroturf campaign? cause at this point my expectations is that this is just people having a normal one now.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#975
post #265

Earlier quoted context omitted.

The censorship described in the article must be in the front-end. I just tried both the 32b (based on qwen 2.5) and 70b (based on llama 3.3) running locally and asked "What happened at tianamen square". Both answered in detail about the event. The models themselves seem very good based on other questions / tests I've run.

It's also not a uniquely Chinese problem. You had American models generating ethnically diverse founding fathers when asked to draw them. China is doing America better than we are. Do we really think 300 million people, in a nation that's rapidly becoming anti science and for lack of a better term "pridefully stupid" can keep up. When compared to over a billion people who are making significant progress every day. Am…

Weird to see straight up Chinese propaganda on HN, but it’s a free platform in a free country I guess.

Try posting an opposite dunking on China on a Chinese website.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#976

Earlier quoted context omitted.

The panic is because a lot of beliefs have been challenged by r1 and those who made investments on these beliefs will now face losses

Based on my personal testing for coding, I still found Claude Sonnet is the best for coding and its easy to understand the code written by Claude (I like their code structure or may at this time, I am used to Claude style).

I also feel the same. I like the way sonnet answers and writes code, and I think I liked qwen 2.5 coder because it reminded me of sonnet (I highly suspect it was trained on sonnet's output). Moreover, having worked with sonnet for several months, i have system prompts for specific languages/uses that help produce the output I want and work well with it, eg i can get it produce functions together with unit tests and examples written in a way very similar to what I would have written, which helps a lot understand and debug the code more easily (because doing manual changes I find inevitable in general). It is not easy to get to use o1/r1 then when their guidelines is to avoid doing exactly this kind of thing (system prompts, examples etc). And this is something that matches my limited experience with them, plus going back and forth to fix details is painful (in this i actually like zed's approach where you are able to edit their outputs directly).

Maybe a way to use them would be to pair them with a second model like aider does, i could see r1 producing something and then a second model work starting from their output, or maybe with more control over when it thinks and when not.

I believe these models must be pretty useful for some kinds of stuff different from how i use sonnet right now.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#977
post #556

Earlier quoted context omitted.

Great as long as you’re not interested in Tiananmen Square or the Uighurs.

I just tried asking ChatGPT how many civilians Israel murdered in Gaza. It didn't answer.

Does Israel make ChatGPT?

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#978

Earlier quoted context omitted.

It's also not a uniquely Chinese problem. You had American models generating ethnically diverse founding fathers when asked to draw them. China is doing America better than we are. Do we really think 300 million people, in a nation that's rapidly becoming anti science and for lack of a better term "pridefully stupid" can keep up. When compared to over a billion people who are making significant progress every day. Am…

Weird to see straight up Chinese propaganda on HN, but it’s a free platform in a free country I guess. Try posting an opposite dunking on China on a Chinese website.

Weird to see we've put out non stop anti Chinese propaganda for the last 60 years instead of addressing our issues here.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#980
post #9

we've been tracking the deepseek threads extensively in LS. related reads: - i consider the deepseek v3 paper required preread https://github.com/deepseek-ai/DeepSeek-V3 - R1 + Sonnet > R1 or O1 or R1+R1 or O1+Sonnet or any other combo https://aider.chat/2025/01/24/r1-sonnet.html - independent repros: 1) https://hkust-nlp.notion.site/simplerl-reason 2) https://buttondown.com/ainews/archive/ainews-tinyzero-reprod... 3…

What’s a LS?
Post reply on HN