Live data from Hacker News

GPT-6 Astra

openai.com

531–540 of 1001 posts

Re: GPT-6 Astra

#531
post #230

I want to take a step back: So, this is GPT-6 -- the natural number version release comparable to GPT-4 and GPT-5 from the past few years. The ARC-AGI-3 score is obviously impressive at 99.9% (we'll need to wait for more details on how they used the response API harness on GPT-6 Astra, wrt reasoning retention and compaction), but every other benchmarks seems to be a relatively modest improvement, comparable with any…

• 98.6% on ARC-AGI-3 • 97.6% on frontier math • 95.9% on CAD • 100% on ExploitBench Nothing modest about it

Except the release announcement. You know, the thing the OP you're responding to is specifically pointing out?

Re: GPT-6 Astra

#532
post #396

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

To each their own. Personally I will start feeling the AGI as soon as we move from chatting about benchmark results to learn that some lab just announced the discovery of tens of novel treatments for rare diseases. Maybe I'm too boring but it seems quite pointless to have this same prediction game every time a new model is released.

Wouldn’t that be ASI? I.e. surpassing humans by outputting novel treatments at a far greater rate than normal humans?

Re: GPT-6 Astra

#535
post #339

It's fun, but every new model release makes me even less interested to create cool stuff. Like, what's the point, if the next AI can do it in 5 seconds?

> Like, what's the point, if the next AI can do it in 5 seconds? Live a life doing whatever makes you happy. Post-work society is an inevitability if we don't destroy our planet.

I think you're assuming work's only function is getting things done. Work also is a crowd control tool.

Re: GPT-6 Astra

#536

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

Their definition of AGI is "when we can't invent any more tests where it fails"

Re: GPT-6 Astra

#539

“allowing non-technical people to create and play custom games that go beyond rudimentary elements” Proceeds to generate the most generic, rudimentary, and unoriginal clone of Mario Kart

It's worse than that, someone else generated it using and then put it on a static page. We just have to take their word for it that GPT6 can do this. It probably can. It's not really an impressive test anymore. Claude Fable can do it. Opus can do it. I've been making one-shotted games with models for a while now, to test out their capabilities, and they all pretty much come out like this - generic bland and basic, using three.js with rudimentary controls and zero gameplay other than collecting points.

Here's a one-shotted submarine game I made with Fable a few weeks back - https://roryok.com/games/deepdive3d.html. One prompt, and I think it's deeper than this (if you'll pardon the pun)

Re: GPT-6 Astra

#540

OpenAI is killing it now that they are more focused. Killing projects like Sora et al have seen it go from irrelevant to level footing with Anthropic. Sol is so much better than Fable 5. Then we get Astra (yet to use it) few days after Fable 5.1 (which is very impressive). Codex is slightly better than Claude Code. Good on Sam Altman getting back to basics and turning OpenAI around.

I think it mostly shows that there is no moat and the only advantage the U.S companies have over the Chinese is more compute. Qwen Max, Kimi K3, GLM 5.3 are really close to Opus/Sol/Fable/Astra and they are open weights.

[flagged]
Post reply on HN