Live data from Hacker News

An analysis of DeepSeek's R1-Zero and R1

arcprize.org

151–160 of 280 posts

Re: An analysis of DeepSeek's R1-Zero and R1

#151

Earlier quoted context omitted.

What's it called when you describe an app with sufficient detail that a computer can carry out the processes you want? Where will the record of those clarifying questions and updates be kept? What if one developer asks the AI to surreptitiously round off pennies and put those pennies into their bank account? Where will that change be recorded, will humans be able to recognize it? What if two developers give it confli…

> What's it called when you describe an app with sufficient detail that a computer can carry out the processes you want? You're wrong here. The entire point is that these are not computers as we used to think of them. These things have common sense; they can analyse a problem including all the implicit aspects, suggest and evaluate different implementation methods, architectures, interfaces. So the right question is:…

Bold of you to imply that GPT asks questions instead of making baseless assumptions every 5 words, even when you explicitly instruct it to ask questions if it doesn't know. When it constantly hallucinates command line arguments and library methods instead of reading the fucking manual.

It's like outsourcing your project to [country where programmers are cheap]. You can't expect quality. Deep down you're actually amazed that the project builds at all. But it doesn't take much to reveal that it's just a facade for a generous serving of spaghetti and bugs.

And refactoring the project into something that won't crumble in 6 months requires more time than just redoing the project from scratch, because the technical debt is obscenely high, because those programmers were awful, and because no one, not even them, understands the code or wants to be the one who has to reverse engineer it.

Except that AI is actually MUCH more expensive!

Re: An analysis of DeepSeek's R1-Zero and R1

#152

Earlier quoted context omitted.

> OpenAI, Meta, AWS, AMD, and others have long attempted to eliminate the Nvidia tax, yet failed. Gemini / Google runs and trains on TPUs. You have no incentive to infer on AMD if you need to buy a massive Nvidia cluster to train.

Google was omitted because they own the hardware and the models, but in retrospect, they represent a proof point nearly as compelling as OpenAI. Thanks for the comment. Google has leading models operating on leading hardware, backed by sophisticated tech talent who could facilitate migrations, yet Google still cannot leap over the CUDA moat and capture meaningful inference market share. Yes, training plays a crucial…

> yet Google still cannot leap over the CUDA moat and capture meaningful inference market share.

It's almost as if being a first-mover is more important than whether or not you use CUDA.

Re: An analysis of DeepSeek's R1-Zero and R1

#153

Earlier quoted context omitted.

Nvidia (NVDA) generates revenue with hardware, but digs moats with software. The CUDA moat is widely unappreciated and misunderstood. Dethroning Nvidia demands more than SOTA hardware. OpenAI, Meta, Google, AWS, AMD, and others have long failed to eliminate the Nvidia tax. Without diving into the gory details, the simple proof is that billions were spent on inference last year by some of the most sophisticated techno…

> OpenAI, Meta, AWS, AMD, and others have long attempted to eliminate the Nvidia tax, yet failed. Gemini / Google runs and trains on TPUs. You have no incentive to infer on AMD if you need to buy a massive Nvidia cluster to train.

Meta trains on Nvidia and infers on AMD. There is incentive if your inference costs are high.

Re: An analysis of DeepSeek's R1-Zero and R1

#154
post #124
post #34

Earlier quoted context omitted.

People talk about Groq and Cerberus as competitors but it seems to me their manufacturing process makes the availability of those chips extremely limited. You can call up Nvidia and order $10B worth of GPUs and have them delivered the next week. Can't say the same for these specialty competitors.

> You can call up Nvidia and order $10B worth of GPUs and have them delivered the next week Nvidia sold $14.5 billion of datacenter hardware in the third quarter of their fiscal 2024 and that led to severe supply constraints, with estimate lead times for H100's up to 52 weeks some places, so no you can't, as that $14.5 billion was clearly capped by their ability to supply, not demand. You're right, though, that Groq…

Groq chips have 230 mb of sram memory. Good luck running 670B model on those chips, even without supply constraints.

Re: An analysis of DeepSeek's R1-Zero and R1

#155

Earlier quoted context omitted.

> What's it called when you describe an app with sufficient detail that a computer can carry out the processes you want? You're wrong here. The entire point is that these are not computers as we used to think of them. These things have common sense; they can analyse a problem including all the implicit aspects, suggest and evaluate different implementation methods, architectures, interfaces. So the right question is:…

Bold of you to imply that GPT asks questions instead of making baseless assumptions every 5 words, even when you explicitly instruct it to ask questions if it doesn't know. When it constantly hallucinates command line arguments and library methods instead of reading the fucking manual. It's like outsourcing your project to [country where programmers are cheap]. You can't expect quality. Deep down you're actually amaz…

Of course, but who's talking about today's tools? They're definitely not able to act like an independent, competent development team. Yet. But if we limit ourselves to the here-and-now, we might be like people talking about GPT3 five years ago: "yes it does spit out a few lines of code, which sometimes even compiles. When it doesn't forget half way and starts talking about unicorns".

We're talking about the tools of tomorrow, which, judging by the extremely rapid progress, I think is only a few (3-5) years away.

Anyway, I had great experiences with Claude and DeepSeek.

Re: An analysis of DeepSeek's R1-Zero and R1

#156
post #20

"The o3 system demonstrates the first practical, general implementation of a computer adapting to novel unseen problems" Yet, they said when it was announced: "OpenAI shared they trained the o3 we tested on 75% of the Public Training set. They have not shared more details. We have not yet tested the ARC-untrained model to understand how much of the performance is due to ARC-AGI data." These two statements are complet…

Your quote is accurate from here:

https://arcprize.org/blog/oai-o3-pub-breakthrough

They were talking about training on the public dataset -- OpenAI tuned the o3 model with 75% of the public dataset. There was some idea/hope that these LLMs would be able to gain enough knowledge in the latent space that they would automatically do well on the ARC-AGI problems. But using 75% of the public training set for tuning puts them at the about same challenge level as all other competitors (who use 100% of training).

In the post they were saying they didn't have a chance to test the o3 model's performance on ARC-AGI "out of-the-box", which is how the 14% scoring R1-zero was tested (no SFT, no search). They have been testing the LLMs out of the box like this to see if they are "smart" wrt the problem set by default.

Re: An analysis of DeepSeek's R1-Zero and R1

#157
post #26

Earlier quoted context omitted.

For inference Nvidia has more significant competition than for training. See Groq, Google's TPU's etc.

Nvidia (NVDA) generates revenue with hardware, but digs moats with software. The CUDA moat is widely unappreciated and misunderstood. Dethroning Nvidia demands more than SOTA hardware. OpenAI, Meta, Google, AWS, AMD, and others have long failed to eliminate the Nvidia tax. Without diving into the gory details, the simple proof is that billions were spent on inference last year by some of the most sophisticated techno…

It's unclear why this drew downvotes, but to reiterate, the comment merely highlights historical facts about the CUDA moat and deliberately refrains from assertions about NVDA's long-term prospects or that the CUDA moat is unbreachable.

With mature models and minimal CUDA dependencies, migration can be justified, but this does not describe most of the LLM inference market today nor in the past.

Re: An analysis of DeepSeek's R1-Zero and R1

#158

Earlier quoted context omitted.

Google was omitted because they own the hardware and the models, but in retrospect, they represent a proof point nearly as compelling as OpenAI. Thanks for the comment. Google has leading models operating on leading hardware, backed by sophisticated tech talent who could facilitate migrations, yet Google still cannot leap over the CUDA moat and capture meaningful inference market share. Yes, training plays a crucial…

> yet Google still cannot leap over the CUDA moat and capture meaningful inference market share. It's almost as if being a first-mover is more important than whether or not you use CUDA.

Both matter quite a bit. The first-mover advantage obviously rewards OEMs in a first-come, first-serve order, but CUDA itself isn't some light switch that OEMs can flick and get working overnight. Everyone would do it if it was easy, and even Google is struggling to find buy-in for their TPU pods and frameworks.

Short-term value has been dependent on how well Nvidia has responded to burgeoning demands. Long-term value is going to be predicated on the number of Nvidia alternatives that exist, and right now the number is still zero.

Re: An analysis of DeepSeek's R1-Zero and R1

#159

> But now with reasoning systems and verifiers, we can create brand new legitimate data to train on. This can either be done offline where the developer pays to create the data or at inference time where the end user pays! > This is a fascinating shift in economics and suggests there could be a runaway power concentrating moment for AI system developers who have the largest number of paying customers. Those customers…

> The most promising idea is to use reasoning models to generate data, and then train our non-reasoning models with the reasoning-embedded data. Why is it promising, aren’t you potentially amplifying AI biases and errors?

it seems to work and seems very scalable, "reasoning" helps to counter biases (answers become longer, ie. the system uses more tokens which means more time to answer a question -- likely longer answers allow better differentiation of answers from each other in the "answer space")

https://newsletter.languagemodels.co/i/155812052/large-scale...

also from the posted article

"""

The R1-Zero training process is capable of creating its own internal domain specific language (“DSL”) in token space via RL optimization.

This makes intuitive sense, as language itself is effectively a reasoning DSL.

"""

Re: An analysis of DeepSeek's R1-Zero and R1

#160

> But now with reasoning systems and verifiers, we can create brand new legitimate data to train on. This can either be done offline where the developer pays to create the data or at inference time where the end user pays! > This is a fascinating shift in economics and suggests there could be a runaway power concentrating moment for AI system developers who have the largest number of paying customers. Those customers…

> The most promising idea is to use reasoning models to generate data, and then train our non-reasoning models with the reasoning-embedded data. DeepSeek did precisely this with their LLama fine-tunes. You can try the 70B one here (might have to sign up): https://groq.com/groqcloud-makes-deepseek-r1-distill-llama-7...

Yes, but I meant it slightly differently than the distills.

The idea is to create the next gen SOTA non reasoning model with synthetic reasoning training data.

Post reply on HN