Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

271–280 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#271
post #180

Earlier quoted context omitted.

Oh, my experience was different. Got the model through ollama. I'm quite impressed how they managed to bake in the censorship. It's actually quite open about it. I guess censorship doesnt have as bad a rep in china as it has here? So it seems to me that's one of the main achievements of this model. Also another finger to anyone who said they can't publish their models cause of ethical reasons. Deepseek demonstrated c…

> I guess censorship doesnt have as bad a rep in china as it has here It's probably disliked, just people know not to talk about it so blatantly due to chilling effects from aforementioned censorship. disclaimer: ignorant American, no clue what i'm talking about.

My guess would be that most Chinese even support the censorship at least to an extent for its stabilizing effect etc.

CCP has quite a high approval rating in China even when it's polled more confidentially.

https://dornsife.usc.edu/news/stories/chinese-communist-part...

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#272

Earlier quoted context omitted.

And (some people here are saying that)* if they are up-to-date is because they're cheating. The copium itt is astounding.

What’s the difference between what they do and what other ai firms do to openai in the us? What is cheating in a business context?

domestically, trade secrets are a thing and you can be sued for corporate espionage. but in an international business context with high geopolitical ramifications? the Soviets copied American tech even when it was inappropriate, to their detriment.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#273

Earlier quoted context omitted.

Apparently the censorship isn't baked-in to the model itself, but rather is overlayed in the public chat interface. If you run it yourself, it is significantly less censored [0] [0] https://thezvi.substack.com/p/on-deepseeks-r1?open=false#%C2...

There's both. With the web interface it clearly has stopwords or similar. If you run it locally and ask about e.g. Tienanmen square, the cultural revolution or Winnie-the-Pooh in China, it gives a canned response to talk about something else, with an empty CoT. But usually if you just ask the question again it starts to output things in the CoT, often with something like "I have to be very sensitive about this subjec…

This is super interesting.

I am not an expert on the training: can you clarify how/when the censorship is "baked" in? Like is the a human supervised dataset and there is a reward for the model conforming to these censored answers?

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#274
post #180

Earlier quoted context omitted.

Apparently the censorship isn't baked-in to the model itself, but rather is overlayed in the public chat interface. If you run it yourself, it is significantly less censored [0] [0] https://thezvi.substack.com/p/on-deepseeks-r1?open=false#%C2...

Oh, my experience was different. Got the model through ollama. I'm quite impressed how they managed to bake in the censorship. It's actually quite open about it. I guess censorship doesnt have as bad a rep in china as it has here? So it seems to me that's one of the main achievements of this model. Also another finger to anyone who said they can't publish their models cause of ethical reasons. Deepseek demonstrated c…

Second this, vanilla 70b running locally fully censored. Could even see in the thought tokens what it didn’t want to talk about.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#275

Earlier quoted context omitted.

I guess all that leetcoding and stack ranking didn't in fact produce "the cream of the crop"...

There's an interesting tweet here from someone who used to work at DeepSeek, which describes their hiring process and culture. No mention of LeetCoding for sure! https://x.com/wzihanw/status/1872826641518395587

they almost certainly ask coding/technical questions. the people doing this work are far beyond being gatekept by leetcode

leetcode is like HN’s “DEI” - something they want to blame everything on

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#276
post #151

Aside from the usual Tiananmen Square censorship, there's also some other propaganda baked-in: https://prnt.sc/HaSc4XZ89skA (from reddit)

Who cares? I ask O1 how to download a YouTube music playlist as a premium subscriber, and it tells me it can't help. Deepseek has no problem.

Do you use the chatgpt website or the api? I suspect these are problems related to the openai's interface itself rather than the models. I have problems getting chatgpt to find me things that it may think it may be illegal or whatever (even if they are not, eg books under CC license). With kagi assistant, with the same openai's models I have not had any such issues. I suspect that should hold in general for api calls.

Also, kagi's deepseek r1 answers the question about about propaganda spending that it is china based on stuff it found on the internet. Well I dont care what the right answer is in any case, what imo matters is that once something is out there open, it is hard to impossible to control for any company or government.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#277
post #12

Earlier quoted context omitted.

almost certainly (see chart) https://www.latent.space/p/reasoning-price-war (disclaimer i made it)

I understand you were trying to make “up and to the right” = “best”, but the inverted x-axis really confused me at first. Not a huge fan. Also, I wonder how you’re calculating costs, because while a 3:1 ratio kind of sort of makes sense for traditional LLMs… it doesn’t really work for “reasoning” models that implicitly use several hundred to several thousand additional output tokens for their reasoning step. It’s alm…

i mean the sheet is public https://docs.google.com/spreadsheets/d/1x9bQVlm7YJ33HVb3AGb9... go fiddle with it yourself but you'll soon see most models hve approx the same input:output token ratio cost (roughly 4) and changing the input:output ratio assumption doesnt affect in the slightest what the overall macro chart trends say because i'm plotting over several OoMs here and your criticisms have the impact of actually the 100:1 ratio starts to trend back toward parity now because of the reasoning tokens, so the truth is somewhere between 3:1 and 100:1.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#279

Earlier quoted context omitted.

Latest GPUs and efficiency are not mutually exclusive, right? If you combine them both presumably you can build even more powerful models.

Not necessarily if you are pushing against a data wall. One could ask: after adjusting for DS efficiency gains how much more compute has OpenAI spent? Is their model correspondingly better? Or even DS could easily afford more than $6 million in compute but why didn't they just push the scaling?

right except that r1 is demoing the path of approach for moving beyond the data wall

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#280

Earlier quoted context omitted.

I would argue there is too little hype given the downloadable models for Deep Seek. There should be alot of hype around this organically. If anything, the other half good fully closed non ChatGPT models are astroturfing. I made a post in december 2023 whining about the non hype for Deep Seek. https://news.ycombinator.com/item?id=38505986

Possible for that to also be true! There’s a lot of astroturfing from a lot of different parties for a few different reasons. Which is all very interesting.

Ye I mean in practice it is impossible to verify. You can kind of smell it though and I smell nothing here, eventhough some of 100 listed authors should be HN users and write in this thread.

Some obvious astroturf posts on HN seem to be on the template "Watch we did boring coorparate SaaS thing X noone cares about!" and then a disappropiate amount of comments and upvotes and 'this is a great idea', 'I used it, it is good' or congratz posts, compared to the usual cynical computer nerd everything sucks especially some minute detail about the CSS of your website mindset you'd expect.

Post reply on HN