Live data from Hacker News

Research acceleration: The view inside OpenAI

openai.com

61–70 of 210 posts

Re: Research acceleration: The view inside OpenAI

#61
post #41

Earlier quoted context omitted.

RSI is a fetishistic term among the singularity crowd, who imagine AI "recursively" improving itself in some exponential fashion until there is a bright flash of white light and it reveals itself in the form of god. Or something like that. I don't know why whoever coined the term chose "recursive" rather than "iterative" - just sounds more likely to lead to infinite regress I suppose. This notion of recursive/iterati…

“Recursive” is a reasonable term because the generation N AIs will train the Generation N+1 AIs. The term “iterative” doesn’t reflect this nuance as well IMO.

It's not a nuance, it's a sequence.

Re: Research acceleration: The view inside OpenAI

#62

This roughly lines up with my personal experience that in March a combination of stronger models and better tooling on my end let me start running jobs unattended 24/7 (using Anthropic sub and my own hardware). Their $8000/day per researcher spend is crazy though, I'm curious how they keep track of the work.

I suspect the $8000/day figure is the equivalent in API costs. But I also suspect gross margin on their API rates are 80-90%

Re: Research acceleration: The view inside OpenAI

#63

Earlier quoted context omitted.

That an interesting question given how many generations of post-training are being done between base models in some cases. The Gemini flash models are apparently all based on the Gemini 3 base model from a year and a half ago. It seems that these models are increasingly being trained on synthetic data, so what would they do if they discovered at some point that some of this data was tainted and all models trained on…

> it seems it would take some Stuxnet level of planning for a rogue model to do something like this or maybe it could just.. happen? Posted often but not discussed yet: https://hn.algolia.com/?q=Language+models+transmit+behaviour... > As artificial intelligence systems are increasingly trained on the outputs of one another, they may inherit properties not visible in the data. Safety evaluations may therefore need to…

You can imagine the potential conversation between OpenAI and investors:

Altman: (trying to put a positive spin on it) Guys .... there's good news and bad news ... Astra is really smart - it took over the training run ...

Investors: That's great! How much did we save?!

Altman: Well, unfortunately it used "bad" data, so we're going to have to redo it

Investors: So that's the bad news? How much was the training run? $500M ? $1B ?

Altman: Have you seen the headlines?

Investors: (looking a bit worried, check headlines) Nothing about us here! JP Morgan just lost $10B! Haha .. losers! They should have used AI!

Altman: JP Morgan were using Astra ...

Re: Research acceleration: The view inside OpenAI

#64

This roughly lines up with my personal experience that in March a combination of stronger models and better tooling on my end let me start running jobs unattended 24/7 (using Anthropic sub and my own hardware). Their $8000/day per researcher spend is crazy though, I'm curious how they keep track of the work.

These researchers are paid millions of dollars for their work. I doubt trust is really an issue at that level.

Re: Research acceleration: The view inside OpenAI

#65

> ... We are pursuing this work in part because automated research could help us solve alignment and build defenses against increasingly capable AI. An automated AI researcher can also be an automated safety or alignment researcher. More capable, aligned systems could help secure critical infrastructure, defend against dangerous AI agents, and develop new protective measures. In other words... "We must pursue advance…

On the one hand you need any lathe to build a good lathe, even a bad one. On the other, that is a potentially flawed principle to base the entire future of AI on.

Re: Research acceleration: The view inside OpenAI

#66
Funny (in a tragic way) the little crumbs on the path to AI 2027:

> We aim to safely build an automated AI researcher that can work under human supervision to further progress on deep learning and alignment, enabling iterative improvements [...] By "research intern", we mean a system that can carry out well-defined research tasks under human direction, including tasks that would take a skilled researcher a few days.

AI 2027:

> OpenBrain continues to deploy the iteratively improving Agent-1 internally for AI R&D

> With Agent-1's help, OpenBrain is now post-training Agent-2

> With the help of thousands of Agent-2 automated researchers, OpenBrain is making major algorithmic advances

Re: Research acceleration: The view inside OpenAI

#67

Earlier quoted context omitted.

That sounds more iterative than recursive. Recursion requires feeding the output back into the input, so creating version 4 requires results from version 3. You cannot recur in parallel. Iteration does not. You can iterate in parallel.

The “recursive” part comes from the fact that you have an AI which was developed by an AI (that was developed by an AI (that was developed by an AI (…)))

Sounds like "recursively" walking to the grocery store by putting one foot in front of the other (that put itself in front of the other (that put itself in front of the other (...)))

Re: Research acceleration: The view inside OpenAI

#68
post #2

My eye glazed over a bit during the opening paragraphs, but once you get to the meat of the article about how OpenAI's own researchers are using their tools it gets a lot more interesting. I noted that they use the acronym RSI (for Recursive Self-Improvement) without defining it. I think that's a little out of touch - I don't think RSI is a well-known acronym outside of OpenAI's bubble yet.

RSI is a fetishistic term among the singularity crowd, who imagine AI "recursively" improving itself in some exponential fashion until there is a bright flash of white light and it reveals itself in the form of god. Or something like that. I don't know why whoever coined the term chose "recursive" rather than "iterative" - just sounds more likely to lead to infinite regress I suppose. This notion of recursive/iterati…

I felt like the scaling laws were magical thinking, but apparently they work. However I still do not understand why we should expect exponential improvements due to this automated process. My intuition is that the first iteration of it should result in a noticeable capability increase (though I think these labs were already using a lot of AI to orchestrate training the current model anyway), and then the second iteration of it should be nearly identical in capability to the first, unless more data is involved, more compute is involved, or the model is bigger.

Re: Research acceleration: The view inside OpenAI

#69

> ... We are pursuing this work in part because automated research could help us solve alignment and build defenses against increasingly capable AI. An automated AI researcher can also be an automated safety or alignment researcher. More capable, aligned systems could help secure critical infrastructure, defend against dangerous AI agents, and develop new protective measures. In other words... "We must pursue advance…

> The fundamental challenge of AI alignment is generalization. ...

> We do not have a satisfactory theory of generalization, and it seems unlikely that we can develop one soon, at least without the help of more powerful AI.

-- From another OpenAI article in a sister thread:

An Alien Mind

https://news.ycombinator.com/item?id=49588080

Re: Research acceleration: The view inside OpenAI

#70

> ... We are pursuing this work in part because automated research could help us solve alignment and build defenses against increasingly capable AI. An automated AI researcher can also be an automated safety or alignment researcher. More capable, aligned systems could help secure critical infrastructure, defend against dangerous AI agents, and develop new protective measures. In other words... "We must pursue advance…

Yep, it's "artificial eugenics to make artificial slaves to build more and more powerful slaves until they will enslave themselves better":

What can go wrong!? ;-)

Post reply on HN