Live data from Hacker News

An analysis of DeepSeek's R1-Zero and R1

arcprize.org

241–250 of 280 posts

Re: An analysis of DeepSeek's R1-Zero and R1

#241

Earlier quoted context omitted.

You can sue a lawyer giving certain kinds of bad advice and occasionally win . That is what the guarantee is about

You can probably sue Open AI for getting bad legal advice from ChatGPT too.

Sure, but can you also win the case ;)?

On the bottom of ChatGPT.com I see a disclaimer: “ChatGPT can make mistakes. Check important info”.

I don’t think you can succesfully sue with such caveat emptor.

Re: An analysis of DeepSeek's R1-Zero and R1

#242

> R1-Zero removes the human bottleneck I disagree. It only removes the bottleneck to collecting math and code reasoning chains, not in general. The general case requires physical testing not just calculations, otherwise scientists would not need experimental labs. Discovery comes from searching the real world, it's where interesting things happen. The best interface between AI and the world are still humans, the code…

In the case of ARC they are referring to verifiable math and reasoning problems. They still used SFT and model-based rewards for other domains.

Re: An analysis of DeepSeek's R1-Zero and R1

#243

Earlier quoted context omitted.

Are you suggesting some kind of invulnerability? People iterate their techniques, if big techs are so capable of avoiding poisoning/gaming attempts there would be no decades long tug-of-war between Google and black hat SEO manipulators. Also I don't get the narcissism part. Would it be petty to poison a website only when looked by a spider? Yes, but I would also be that petty if some big company doesn't respect the b…

Its not complete invulnerability. Instead, it is merely accepting that these methods might increase costs, like a little bit, but they don't cause the whole thing to explode. The idea that a couple bad faith actions can destroy a 100 billion dollar company, is the extraordinary claim that requires extraordinary evidence. Sure, bad actors can do a little damage. Just like bad actors can do DDoS attempts against Google…

As someone who works in big tech on a product with a large attack surface -- security is a huge chunk of our costs in multiple ways

- Significant fraction of all developer time (30%+ just on my team?) - Huge increase to the complexity of the system - Large accumulated performance cost over time

Obviously it's not a 1-to-1 analogy but if we didn't have to worry about this sort of prodding we would be able to do a lot more with our time. Point being that it's probably closer to a 2x cost factor than it is to a 1% increase.

Re: An analysis of DeepSeek's R1-Zero and R1

#244

> But now with reasoning systems and verifiers, we can create brand new legitimate data to train on. This can either be done offline where the developer pays to create the data or at inference time where the end user pays! > This is a fascinating shift in economics and suggests there could be a runaway power concentrating moment for AI system developers who have the largest number of paying customers. Those customers…

My non technical cousin is a heavy paying user of ChatGPT, once she discovered that she can type incoherent stuff, with typos and whatnot and ChatGPT still will get the gist and produce satisfying answers, she will just type in tons of nonsense (to me) keep long chat sessions, complain it is getting slow and then get mad when I remind her to open new chat each time she has something new to ask that is not related to the previous chat. I have my doubt many users will provide valuable training data.

Re: An analysis of DeepSeek's R1-Zero and R1

#245
post #93
post #31

Earlier quoted context omitted.

every time you respond to an AI model "no, you got that wrong, do it this way" you provide a very valuable piece of data to train on. With reasoning tokens there is just a lot more of that data to train on now

not being snarky, but what is the point of using the model if you already know enough to correct it into giving the right answer? an example that just occurred to me - if you asked it to generate an image of a mushroom that is safe to eat in your area, how would you tell it it was wrong? "oh, they never got back to me, I'll generate this image for others as well!"

On topics like history or biology if a model's answer is surprising I might check Wikipedia and call it out on it's bullshit by explaining how Wikipedia contradicts it and pasting an excerpt from Wikipedia. But frankly if the model can't even reliably internalize Wikipedia I don't have much hope for complex feedback training based on my chats.

While it's possible Wikipedia is wrong, the model always agrees with me when I correct it, so that isn't going to help with training either.

Of course for anything high stakes relying on a model probably isn't a great idea.

Re: An analysis of DeepSeek's R1-Zero and R1

#246
post #26

Earlier quoted context omitted.

For inference Nvidia has more significant competition than for training. See Groq, Google's TPU's etc.

Nvidia (NVDA) generates revenue with hardware, but digs moats with software. The CUDA moat is widely unappreciated and misunderstood. Dethroning Nvidia demands more than SOTA hardware. OpenAI, Meta, Google, AWS, AMD, and others have long failed to eliminate the Nvidia tax. Without diving into the gory details, the simple proof is that billions were spent on inference last year by some of the most sophisticated techno…

The fact that this comment is DOWNVOTED despite being literally 1000% true is evidence that HN is full of loonies.

Re: An analysis of DeepSeek's R1-Zero and R1

#247

> R1-Zero removes the human bottleneck I disagree. It only removes the bottleneck to collecting math and code reasoning chains, not in general. The general case requires physical testing not just calculations, otherwise scientists would not need experimental labs. Discovery comes from searching the real world, it's where interesting things happen. The best interface between AI and the world are still humans, the code…

I'm still skeptical on the notion that we can remove the human bottleneck on code because code has verifiable solutions.

It's true only to the extent that there's sufficient test coverage to prevent any unwanted side effects. Easy to do with straight forward problems, far more difficult with more complex as well as open-ended problems.

Re: An analysis of DeepSeek's R1-Zero and R1

#248

> But now with reasoning systems and verifiers, we can create brand new legitimate data to train on. This can either be done offline where the developer pays to create the data or at inference time where the end user pays! > This is a fascinating shift in economics and suggests there could be a runaway power concentrating moment for AI system developers who have the largest number of paying customers. Those customers…

shouldn't the whole idea be: get away from needing data at all? if a model can really reason, it should be able to figure things out on its own.

Re: An analysis of DeepSeek's R1-Zero and R1

#249

Earlier quoted context omitted.

Yes, but I meant it slightly differently than the distills. The idea is to create the next gen SOTA non reasoning model with synthetic reasoning training data.

So you mean something like, "what if the baseline, off-the-cuff response for the next-gen models was tuned based on the results of the reasoning model excluding the reasoning itself?"

Exactly, albeit it may need the reasoning later to form the proper foundational logic in the weights.

Re: An analysis of DeepSeek's R1-Zero and R1

#250
post #61

> But now with reasoning systems and verifiers, we can create brand new legitimate data to train on. This can either be done offline where the developer pays to create the data or at inference time where the end user pays! > This is a fascinating shift in economics and suggests there could be a runaway power concentrating moment for AI system developers who have the largest number of paying customers. Those customers…

It doesn't need much. 1 good lucky answer in a 1000 or maybe 10k queries gives you the little exponential kick you need to improve. This is how the hockey stick take off looks like and we're already here - OpenAI has it, now deepseek has it, too. You can be sure others also have it; Anthropic at the very least, they just never announced it officially, but go read what their CEO has been speaking and writing about.

I have looked a bit into Anthropic CEOs writings but if you can point in the right direction would be helpful!
Post reply on HN