Live data from Hacker News

Research acceleration: The view inside OpenAI

openai.com

191–200 of 210 posts

Re: Research acceleration: The view inside OpenAI

#191

Funny (in a tragic way) the little crumbs on the path to AI 2027: > We aim to safely build an automated AI researcher that can work under human supervision to further progress on deep learning and alignment, enabling iterative improvements [...] By "research intern", we mean a system that can carry out well-defined research tasks under human direction, including tasks that would take a skilled researcher a few days.…

This year has really cemented Daniel Kokotajlo‘s reputation for me. Even if the rest of the predictions are way off from this point on, its really impressive how accurate his forecast for 2026 has been

a link to article he is referring to.

https://ai-2027.com/

worth reading

Re: Research acceleration: The view inside OpenAI

#192
post #161

Earlier quoted context omitted.

I would actually like to see them solve these problems, I don't care who comes up with solutions to curing cancer, etc

I think people very much should care about who ends up owning these solutions. The person or entity that controls things like that just has more power, which isn’t necessarily a good thing

It is a good thing when it didn't exist before and it does exist now and wouldn't have existed without them.

If they profit immensely from curing cancer, good.

Re: Research acceleration: The view inside OpenAI

#193

Earlier quoted context omitted.

Yes, you need some kind of other source of truth. I think the best way to get that is to do clean room development with a different agent, but ultimately if you give them the wrong idea they'll do the wrong thing. The other thing I do, not as much as I should, but it's very powerful, is to generate spikes and deliberately throw them away to understand how to prompt better. Like I generated a swift version of the reac…

If you've managed people, these are all familiar problems. I found you need much more than a functional specification, you also need motivation, background, related work, ideas tried, etc., because those help disambiguate the right path in the inevitable situation where your original task description is unclear or conflicts with itself.

Yes I'm using the entire consultancy stack - define values, etc, and work your way down the "where do these not match reality on the ground and need change", but for little robot people instead of (arguably less messy) humans.

Re: Research acceleration: The view inside OpenAI

#194

Earlier quoted context omitted.

My only experience in >24h agents is with economically sane models (one of GLM5.2, 5.3-flash for orchestration, DSV4-flash for implementation, and glm5.3|sol|kimi3 agents + subagents reviewing at the end) Over 24h my token spend is I'm not sure what the point would be though unless working on some kind of optimization problem -- it takes me days to review <24h of the agent's output. It's almost always near enough to…

This sounds like more work than just writing the code yourself. You'll say it isn't. I don't believe you.

That is an incorrect presumption - I think it's plausible this is more work; it's certainly far more taxing.

I'm at a point in my career where a small minority of my time is coding. The AIs can do in a day what would have taken me a week uninterrupted with acceptable (in some cases inferior prior to human feedback--but in some cases superior!) quality.

As I do not have 10 let alone 40 hours per week to devote to coding I think it increases the amount of high quality work product I can create with a given time investment. As I review it I merge small independent units and decompose the work.

All that is to say I don't really like it - but I suspect for most *well defined* coding tasks human produced code from highly experienced engineers will largely cease to exist in the next year -- getting cheap/relatively horrible models to produce good code is now straightforward.

OTOH I never use AI for any human facing communication outside of making my writing shorter. IMO AI slop "documents" are almost certainly a drag on organizational productivity.

Re: Research acceleration: The view inside OpenAI

#195

Earlier quoted context omitted.

My only experience in >24h agents is with economically sane models (one of GLM5.2, 5.3-flash for orchestration, DSV4-flash for implementation, and glm5.3|sol|kimi3 agents + subagents reviewing at the end) Over 24h my token spend is I'm not sure what the point would be though unless working on some kind of optimization problem -- it takes me days to review <24h of the agent's output. It's almost always near enough to…

It only really makes sense for problems that are complex and require iterations that don't themselves require much review. E.g. if you want find, PoC, and patch bugs, the output can be reviewed without reading all the traces. Or if you want to write a custom tool that does some job using local LLMs, assembling that pipeline, tuning the prompts, etc takes a long time but reading the final tests + eval data + code is e…

>don't themselves require much review

I'm still wary of any unreviewed code - though my area of work is not tolerant of defects.

Agree on targets / verifiable indications of progress or success being a prerequisite for this being useful - although that covers quite a lot of SWE work.

Re: Research acceleration: The view inside OpenAI

#196
post #180

Earlier quoted context omitted.

"Destroying the ecosystem" is just FUD. And if you don't find "average AI research intern" impressive, I'm not sure what to tell you. Have the goalposts moved so far that open ended problem solving at "average CS student fresh out of the uni" levels is suddenly trivial? Think of what AI was capable of in 2016. Or even 2022. Compare that to now. We had more AI progress in the last five years than I expected to happen…

The ecosystem absolutely is being destroyed. We're looking at anywhere from 3-5C warming by 2060, which is going to be devastating if not flat out apocalyptic. And wherever we're at in 2060 it's not like it's going to stop there, nor is it going to be comfortable until then. Things may start to crumble much sooner. I don't believe AI and data centers have played that much of a role in this though, we could have power…

We're not even looking at 5C of warming by 2100 realistically. Like, that was considered to be an unlikely extreme scenario in 2014 AR5, and also in the tightened down 2021 AR6, and things have happened since! Renewables are cheaper than ever, and Ukrainian war and Iranian war both curbed the appetite for long term fossil fuel power investment.

The median is what, a bit under 3C by 2100? Not even by 2060 - by 2100. And we're in 2026, so that's more than twice as slow as your expectation.

Agreed on AI not being a meaningful factor in climate change though. We'd have to go full "humankind is obsolete" technological singularity to have AI dominate energy use to this extent, and current numbers are nowhere near that. It's a FUD distraction from the real culprits: the fossil fuel energy complex. That's currently lobbying to slow the inevitable energy transition.

Re: Research acceleration: The view inside OpenAI

#197
post #12

Earlier quoted context omitted.

Opus was trained based on it's internal CoT due to a bug for generations. Gemini's depression extended through models. OpenAI has killed people. We've already seen cross gen misalingment.

Source?

OpenAI helped out a mass shooter in Tumbler Ridge, plus all the suicides

Anthropic training on CoT for multi gens: https://www.lesswrong.com/posts/K8FxfK9GmJfiAhgcT/anthropic-...

Can't find anything specifically about the Gemini issue being a training data contamination, but the depression was real:

https://www.businessinsider.com/gemini-self-loathing-i-am-a-...

I think the Gemini depression being persisted across gens via training data was a HN comment I can't find anymore, no strong source.

Re: Research acceleration: The view inside OpenAI

#198
post #180

Earlier quoted context omitted.

The ecosystem absolutely is being destroyed. We're looking at anywhere from 3-5C warming by 2060, which is going to be devastating if not flat out apocalyptic. And wherever we're at in 2060 it's not like it's going to stop there, nor is it going to be comfortable until then. Things may start to crumble much sooner. I don't believe AI and data centers have played that much of a role in this though, we could have power…

We're not even looking at 5C of warming by 2100 realistically. Like, that was considered to be an unlikely extreme scenario in 2014 AR5, and also in the tightened down 2021 AR6, and things have happened since! Renewables are cheaper than ever, and Ukrainian war and Iranian war both curbed the appetite for long term fossil fuel power investment. The median is what, a bit under 3C by 2100? Not even by 2060 - by 2100. A…

> that was considered to be an unlikely extreme scenario

By the same people who just realized we're missing 1.5C as we're blazing past it at mach 12, still accelerating not slowing down? You really believe those guys?

You have to understand that there are several camps of climate science. The mainstream ones like IPCC and UN etc are heavily politicized, they can't publish anything that isn't sugarcoated beyond recognition. At least I assume that's why they're so obviously wrong.

Here's a judgement I think is more realistic

> There is a strong probability that the ambition gap will lead to a temperature rise of 2 to 5 degrees Centigrade compared to pre-industrial temperatures by 2100, the realisation gap to a further rise of several degrees Centigrade.[1,2,VI] There is a danger that the mean temperature will already have risen by 3 degrees Centigrade by 2050.

https://www.dpg-physik.de/veroeffentlichungen/publikationen/...

Re: Research acceleration: The view inside OpenAI

#199
post #198

Earlier quoted context omitted.

We're not even looking at 5C of warming by 2100 realistically. Like, that was considered to be an unlikely extreme scenario in 2014 AR5, and also in the tightened down 2021 AR6, and things have happened since! Renewables are cheaper than ever, and Ukrainian war and Iranian war both curbed the appetite for long term fossil fuel power investment. The median is what, a bit under 3C by 2100? Not even by 2060 - by 2100. A…

> that was considered to be an unlikely extreme scenario By the same people who just realized we're missing 1.5C as we're blazing past it at mach 12, still accelerating not slowing down? You really believe those guys? You have to understand that there are several camps of climate science. The mainstream ones like IPCC and UN etc are heavily politicized, they can't publish anything that isn't sugarcoated beyond recogn…

Yes, I do. IPCC's reports are sensible. They're not unreliable just because they don't support the "doom and burning land" narratives.

By the way, there is no "just realized we're missing 1.5C". That projection was always the very low end of possibilities - the "assume rapid, radical climate action on global level" scenario.

Yes, that's a dumb thing to assume. We've never been on track for it. But the "assume extremely high emissions and no green transition ever, 5C+ by 2100" scenario on the other end is about as unlikely to materialize. Those are the boundaries of the expectation range - not median expectations.

Re: Research acceleration: The view inside OpenAI

#200
post #198

Earlier quoted context omitted.

> that was considered to be an unlikely extreme scenario By the same people who just realized we're missing 1.5C as we're blazing past it at mach 12, still accelerating not slowing down? You really believe those guys? You have to understand that there are several camps of climate science. The mainstream ones like IPCC and UN etc are heavily politicized, they can't publish anything that isn't sugarcoated beyond recogn…

Yes, I do. IPCC's reports are sensible. They're not unreliable just because they don't support the "doom and burning land" narratives. By the way, there is no "just realized we're missing 1.5C". That projection was always the very low end of possibilities - the "assume rapid, radical climate action on global level" scenario. Yes, that's a dumb thing to assume. We've never been on track for it. But the "assume extreme…

I'm pessimistic. We're still producing more CO2 each year than the last, and several feedback loops are kicking in that accelerate the warming further such as permafrost thawing, arctic and antarctic sea ice disappearing, glaciers are melting, the amazon is being demolished, etc. I don't know how much of this is included in the projections.

I hope the optimists are right, it just doesn't look like it to me at all. It looks to me like we're speeding along right into the worst predictions and beyond. We're building lots of green energy production but it seems to just come in on top of existing and new fossil production not replace it.

It does seem like the CO2 output is plateauing which is good, but we really need it to start declining drastically very soon and I don't really see that happening with the current political climate. Also remember CO2 is far from the only greenhouse gas - methane, nitrous oxide and fluorinated gas emissions all seem to be rising rapidly still.

Post reply on HN