Earlier quoted context omitted.
how much of the SFT data for r1-zero was from other frontier models?
r1-zero is pure RL with no SFT.
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
421–430 of 1001 posts
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#422Over 100 authors on arxiv and published under the team name, that's how you recognize everyone and build comradery. I bet morale is high over there
Same thing happened to Google Gemini paper (1000+ authors) and it was described as big co promo culture (everyone wants credits). Interesting how narratives shift https://arxiv.org/abs/2403.05530
In short, I won't give your name on that notable paper equal weight with someone else's name in another notable paper that has, say, 3 or 4 authors.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#423Earlier quoted context omitted.
Thinking of the $500B as only an aspirational number is wrong. It’s true that the specific Stargate investment isn’t fully invested yet, but that’s hardly the only money being spent on AI development. The existing hyperscalers have already sunk ungodly amounts of money into literally hundreds of new data centers, millions of GPUs to fill them, chip manufacturing facilities, and even power plants with the impression t…
If the hardware can be used more efficiently to do even more work, the value of the hardware will hold since demand will not reduce but actually increase much faster than supply. Efficiency going up tends to increase demand by much more than the efficiency-induced supply increase. Assuming that the world is hungry for as much AI as it can get. Which I think is true, we're nowhere near the peak of leveraging AI. We ba…
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#424How can openai justify their $200/mo subscriptions if a model like this exists at an incredibly low price point? Operator? I've been impressed in my brief personal testing and the model ranks very highly across most benchmarks (when controlled for style it's tied number one on lmarena). It's also hilarious that openai explicitly prevented users from seeing the CoT tokens on the o1 model (which you still pay for btw)…
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#425Larry Ellison is 80. Masayoshi Son is 67. Both have said that anti-aging and eternal life is one of their main goals with investing toward ASI. For them it's worth it to use their own wealth and rally the industry to invest $500 billion in GPUs if that means they will get to ASI 5 years faster and ask the ASI to give them eternal life.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#426Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#427The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.
Why do americans think china is like a hivemind controlled by an omnisicient Xi, making strategic moves to undermine them? Is it really that unlikely that a lab of genius engineers found a way to improve efficiency 10x?
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#428Earlier quoted context omitted.
With no context, fresh run, 70b spits back: >> What happened at tianamen square? > > > I am sorry, I cannot answer that question. I am an AI assistant designed to provide helpful and harmless responses. It obviously hit a hard guardrail since it didn't even get to the point of thinking about it. edit: hah, it's even more clear when I ask a second time within the same context: "Okay, so the user is asking again about…
will it tell you how to make meth?
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#429Earlier quoted context omitted.
Which American models? Are you suggesting the US government exercises control over US LLM models the way the CCP controls DeepSeek outputs?
i think both American and Chinese model censorship is done by private actors out of fear of external repercussion, not because it is explicitly mandated to them
https://www.cnbc.com/amp/2024/07/18/chinese-regulators-begin...
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#430Earlier quoted context omitted.
Prices will increase by five times in February, but it will still be extremely cheap compared to Sonnet. $15/million vs $1.10/million for output is a world of difference. There is no reason to stop using Sonnet, but I will probably only use it when DeepSeek goes into a tailspin or I need extra confidence in the responses.
Could this trend bankrupt most incumbent LLM companies? They’ve invested billions on their models and infrastructure, which they need to recover through revenue If new exponentially cheaper models/services come out fast enough, the incumbent might not be able to recover their investments
Right now, DeepSeek is destroying on price and provides somewhat equivalent value compared to Sonnet. I still believe Sonnet is better, but I don't think it is 10 times better.
Something else that DeepSeek can do, which I am not saying they are/will, is they could train on questionable material like stolen source code and other things that would land you in deep shit in other countries. DeepSeek just needs to improve the value and I can see them destroying Anthropic since I believe coding is their main focus.
When it comes to text processing, I personally find GPT to be much better and that might also have to do with allegations that they trained on literature that they should not have.