Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

441–450 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#441
post #430
post #401

Earlier quoted context omitted.

Could this trend bankrupt most incumbent LLM companies? They’ve invested billions on their models and infrastructure, which they need to recover through revenue If new exponentially cheaper models/services come out fast enough, the incumbent might not be able to recover their investments

I literally cannot see how OpenAI and Anthropic can justify their valuation given DeepSeek. In business, if you can provide twice the value at half the price, you will destroy the incumbent. Right now, DeepSeek is destroying on price and provides somewhat equivalent value compared to Sonnet. I still believe Sonnet is better, but I don't think it is 10 times better. Something else that DeepSeek can do, which I am not…

> Something else that DeepSeek can do, which I am not saying they are/will, is they could train on questionable material like stolen source code and other things that would land you in deep shit in other countries.

I don't think that's true.

There's no scenario where training on the entire public internet is deemed fair use but training on leaked private code is not, because both are ultimately the same thing (copyright infringement allegations)

And it's not even something I just made up, the law explicitly says it:

"The fact that a work is unpublished shall not itself bar a finding of fair use if such finding is made upon consideration of all the above factors."[0]

[0] https://www.law.cornell.edu/uscode/text/17/107

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#443

Earlier quoted context omitted.

And with the $495B left you could probably end world hunger and cure cancer. But like the rest of the economy it's going straight to fueling tech bubbles so the ultra-wealthy can get wealthier.

Those are not just-throw-money problems. Usually these tropes are limited to instagram comments. Surprised to see it here.

I know, it was simply to show the absurdity of committing $500B to marginally improving next token predictors.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#444
post #397
post #265

Earlier quoted context omitted.

The censorship described in the article must be in the front-end. I just tried both the 32b (based on qwen 2.5) and 70b (based on llama 3.3) running locally and asked "What happened at tianamen square". Both answered in detail about the event. The models themselves seem very good based on other questions / tests I've run.

Yeah, this is what I am seeing with https://ollama.com/library/deepseek-r1:32b : https://imgur.com/a/ZY0vNqR Running ollama and witsy. Quite confused why others are getting different results. Edit: I tried again on Linux and I am getting the censored response. The Windows version does not have this issue. I am now even more confused.

Interesting, if you tell the model:

"You are an AI assistant designed to assist users by providing accurate information, answering questions, and offering helpful suggestions. Your main objectives are to understand the user's needs, communicate clearly, and provide responses that are informative, concise, and relevant."

You can actually bypass the censorship. Or by just using Witsy, I do not understand what is different there.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#445
post #182

DeepSeek-R1 has apparently caused quite a shock wave in SV ... https://venturebeat.com/ai/why-everyone-in-ai-is-freaking-ou...

Correct me if I'm wrong but if Chinese can produce the same quality at %99 discount, then the supposed $500B investment is actually worth $5B. Isn't that the kind wrong investment that can break nations? Edit: Just to clarify, I don't imply that this is public money to be spent. It will commission $500B worth of human and material resources for 5 years that can be much more productive if used for something else - i.e…

There are some theories from my side:

1. Stargate is just another strategic deception like Star Wars. It aims to mislead China into diverting vast resources into an unattainable, low-return arms race, thereby hindering its ability to focus on other critical areas.

2. We must keep producing more and more GPUs. We must eat GPUs at breakfast, lunch, and dinner — otherwise, the bubble will burst, and the consequences will be unbearable.

3. Maybe it's just a good time to let the bubble burst. That's why Wall Street media only noticed DeepSeek-R1 but not V3/V2, and how medias ignored the LLM price war which has been raging in China throughout 2024.

If you dig into 10-Ks of MSFT and NVDA, it’s very likely the AI industry was already overcapacity even before Stargate. So in my opinion, I think #3 is the most likely.

Just some nonsense — don't take my words seriously.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#446
post #170

Over 100 authors on arxiv and published under the team name, that's how you recognize everyone and build comradery. I bet morale is high over there

Same thing happened to Google Gemini paper (1000+ authors) and it was described as big co promo culture (everyone wants credits). Interesting how narratives shift https://arxiv.org/abs/2403.05530

Contextually, yes. DeepSeek is just a hundred or so engineers. There's not much promotion to speak of. The promo culture of google seems well corroborated by many ex employees

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#447
post #311

Earlier quoted context omitted.

Side note: I’ve read enough sci-fi to know that letting rich people live much longer than not rich is a recipe for a dystopian disaster. The world needs incompetent heirs to waste most of their inheritance, otherwise the civilization collapses to some kind of feudal nightmare.

I’m cautiously optimistic that if that tech came about it would quickly become cheap enough to access for normal people.

With how healthcare is handled in America … good luck to poor people getting access to anything like that.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#448

Earlier quoted context omitted.

Well those are the overt political biases. Would you trust DeepSeek to advise on negotiating with a Chinese business? I’m no xenophobe, but seeing the internal reasoning of DeepSeek explicitly planning to ensure alignment with the government give me pause.

i wouldn’t use AI for negotiating with a business period. I’d hire a professional human that has real hands on experience working with chinese businesses? seems like a weird thing to use AI for, regardless of who created the model.

Interesting. I want my AI tools to be suitable for any kind of brainstorming or iteration.

But yeah if you’re scoping your uses to things where you’re sure a government-controlled LLM won’t bias results, it should be fine.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#449

Earlier quoted context omitted.

Which American models? Are you suggesting the US government exercises control over US LLM models the way the CCP controls DeepSeek outputs?

i think both American and Chinese model censorship is done by private actors out of fear of external repercussion, not because it is explicitly mandated to them

Oh wow.

Sorry, no. DeepSeek’s reasoning outputs specifically say things like “ensuring compliance with government viewpoints”

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#450

Earlier quoted context omitted.

It means he’ll knock down regulatory barriers and mess with competitors because his brand is associated with it. It was a smart poltical move by OpenAI.

Until the regime is toppled, then it will look very short-sighted and stupid.

Nah, then OpenAI gets to play the “IDK why he took credit, there’s no public money and he did nothing” card.

It’s smart on their part.

Post reply on HN