Live data from Hacker News

An analysis of DeepSeek's R1-Zero and R1

arcprize.org

251–260 of 280 posts

Re: An analysis of DeepSeek's R1-Zero and R1

#251

> But now with reasoning systems and verifiers, we can create brand new legitimate data to train on. This can either be done offline where the developer pays to create the data or at inference time where the end user pays! > This is a fascinating shift in economics and suggests there could be a runaway power concentrating moment for AI system developers who have the largest number of paying customers. Those customers…

shouldn't the whole idea be: get away from needing data at all? if a model can really reason, it should be able to figure things out on its own.

An LLM is just a really good parser connected to a lossy compressed corpus of data.

They need to be open ended and self training to be truly useful.

Reasoning is way far away...

Re: An analysis of DeepSeek's R1-Zero and R1

#252

I think deepseek accidentally also killed google for me, not just chatgpt. Because of the visible reasoning part.

From what I read elsewhere (random reddit comment), the visible reasoning is just "for show" and isn't the process deepseek used to arrive at the result. But if the reasoning has value, I guess it doesn't matter even if it's fake.

There is no way this is true. It is just an example of why Reddit is a fucking joke that you should never read.

I have seen it infer incredibly obscure things in the chain of thought that I was impressed it could piece together.

It is an incredible tool. I would trust it 1000% more than a random person on reddit.

Re: An analysis of DeepSeek's R1-Zero and R1

#253
post #20

"The o3 system demonstrates the first practical, general implementation of a computer adapting to novel unseen problems" Yet, they said when it was announced: "OpenAI shared they trained the o3 we tested on 75% of the Public Training set. They have not shared more details. We have not yet tested the ARC-untrained model to understand how much of the performance is due to ARC-AGI data." These two statements are complet…

They are testing with a different dataset. The authors saying that they have not tested on the version of o3 that has not seen the training set.

Yeah...the whole point is that you're testing the model on something it hasn't seen already. If the problems were in the training set by definition the model has seen them before.

Re: An analysis of DeepSeek's R1-Zero and R1

#254
post #65

Earlier quoted context omitted.

Have you tried suno.ai?

Have _you_? It lost its novelty after a couple of days.

I probably listen to Suno (both my own songs, and songs other people have created) about as often as I listen to Spotify, these days.

Re: An analysis of DeepSeek's R1-Zero and R1

#255
post #46

Earlier quoted context omitted.

This assumes that you give honest feedback. Efforts to feed deployed AI models various epistemic poisons abound in the wild.

> This assumes that you give honest feedback. You don't need honest user feedback because you could judge any message part of a conversation using hindsight. Just ask a LLM to judge if a response is useful, while seeing what messages come after it. The judge model has privileged information. Maybe 5 messages later it turns out what the LLM replied was not a good idea. You can also use related conversations by the sam…

It goes then in the line of https://xkcd.com/810/

Re: An analysis of DeepSeek's R1-Zero and R1

#256

Earlier quoted context omitted.

Its not complete invulnerability. Instead, it is merely accepting that these methods might increase costs, like a little bit, but they don't cause the whole thing to explode. The idea that a couple bad faith actions can destroy a 100 billion dollar company, is the extraordinary claim that requires extraordinary evidence. Sure, bad actors can do a little damage. Just like bad actors can do DDoS attempts against Google…

Grandparent mentioned "we", I guess they refer to a full class of "black hats" avoiding bad faith scraping that eventually could amass to a relatively effective volume of poisoned sites and/or feedback to the model. Obviously a singular poisoned site will never make a difference in a dataset of billions and billions of tokens, much less destroy a 100bn company. That's a straw man, and I think people arguing about poi…

Global coordination for lulz exists, it's called "memes".

Remember Dogecoin or Gamestop; the lulz-oriented meme outbursts had a real impact.

Equally, a particular way to gaslight LLM scrapers may become popular and widespread without any enforcement.

Re: An analysis of DeepSeek's R1-Zero and R1

#257

I predict that the future of LLM's when it comes to coding and software creation is in "custom individually tailored apps". Imagine telling an AI agent what app you want, the requirements and all that and it just builds everything needed from backend to frontend, asks for your input on how things should work, clarifying questions etc. It tests the software by compiling and running it reading errors and failed tests a…

Marvin Minsky promised that an AI would have a PhD, by 1950, and 1960... we are no closer. sorry. We are faster, much faster, 100,000,000 times faster, by we are no closer.

Re: An analysis of DeepSeek's R1-Zero and R1

#258
post #256

Earlier quoted context omitted.

Grandparent mentioned "we", I guess they refer to a full class of "black hats" avoiding bad faith scraping that eventually could amass to a relatively effective volume of poisoned sites and/or feedback to the model. Obviously a singular poisoned site will never make a difference in a dataset of billions and billions of tokens, much less destroy a 100bn company. That's a straw man, and I think people arguing about poi…

Global coordination for lulz exists, it's called "memes". Remember Dogecoin or Gamestop; the lulz-oriented meme outbursts had a real impact. Equally, a particular way to gaslight LLM scrapers may become popular and widespread without any enforcement.

Didn't think of it that way, but I think you're right. As long as memes exist one could argue the LLMs are going to be poisoned in one way or another.

Re: An analysis of DeepSeek's R1-Zero and R1

#259
post #79

Earlier quoted context omitted.

That's going to be much slower and more expensive than writing tests because image/video processing is slower and more expensive than writing tests. And because of lag in using the UI (and re-building the whole application from scratch after every change to test again).

Hm, what if instead of using video of the application… Ok, so if one can have one program snoop on all the rendering calls made by another program, maybe there could be a way of training a common representation of “an image of an application” and “the rendering calls that are made when producing a frame of the display for the application”? Hopefully in a way that would be significantly smaller than the full image dat…

Yes, it would be significantly smaller, but it would look very different depending on your platform, GPU, driver version, etc. -- the model would essentially need to learn how to map "graphics APIs" (e.g. OpenGL, Vulkan, Metal, ...) to "render result" for every combination of API, driver version, and GPU, which I imagine would constitute a significant amount of overhead.

Re: An analysis of DeepSeek's R1-Zero and R1

#260

Earlier quoted context omitted.

If you want to pick apart my hastily concocted examples, well, have fun I guess. My overall point is that ensuring data quality is something OpenAI is probably very good at. They likely have many clever techniques, some of which we could guess at, some of which would surprise us, all of which they’ve validated through extensive testing including with adversarial data. If people want to keep playing pretend that their…

I'm interested in why you think OpenAI is probably very good at ensuring data quality. Also interested if you are trying to troll the resistance into revealing their working techniques.

They buy it through scale ai
Post reply on HN