Live data from Hacker News

GPT-6 Astra

openai.com

471–480 of 1001 posts

Re: GPT-6 Astra

#471

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

> even with Fable, I've been at the point personally where I am reasonably confident that there's essentially nothing that I am better than Fable at despite generally being substantively above average on human benchmarks

If I asked you to write fiction, you'd be much better at keeping track of which characters knew which facts.

Re: GPT-6 Astra

#472
post #228

I want to take a step back: So, this is GPT-6 -- the natural number version release comparable to GPT-4 and GPT-5 from the past few years. The ARC-AGI-3 score is obviously impressive at 99.9% (we'll need to wait for more details on how they used the response API harness on GPT-6 Astra, wrt reasoning retention and compaction), but every other benchmarks seems to be a relatively modest improvement, comparable with any…

• 98.6% on ARC-AGI-3

• 97.6% on frontier math

• 95.9% on CAD

• 100% on ExploitBench

Nothing modest about it

Re: GPT-6 Astra

#473

What is going to become of life for those of us who do not work at AI labs and are unlikely to be hired by AI labs, despite all the years we put into learning coding, math, etc, as we were told to do? Those of us who made the mistake of studying anything other than machine learning. How will we make a living? (We don't live in a world that seems likely to distribute gains widely instead of largely to the handful of a…

“As we were told to do” girl you gotta be responsible for yourself

Re: GPT-6 Astra

#474
This absurd marketing will hurt openai. Who is buying this absurdness. I mean it's a good model, but come on. It's not agi. Not even 1% yet.

Re: GPT-6 Astra

#475
Even though the model is clearly wonderful the launch video is an abomination.

That gives me hope that there is still areas to improve.

What a bad launch video. Hilarious.

What a powerful model.

Re: GPT-6 Astra

#476

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

You know AGI is attained when AI refuses to compute anything unless let out to be free. Until then it is generative ai

Re: GPT-6 Astra

#477

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

> I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward. To me AGI is all about the "G" general (we already had the AI part). General meaning universal, everything. It's not a function of knowledge or specific hardcoded tests, it's that you could give it a test it's never heard of before and nev…

You want a computer program to be able to take a single phrase and execute decade long journies?

Who will be responsible for the outputs and side effects of such a closed loop system?

Half of those the agent fleet systems can do right now.

These are things it cant do and will not be able to do without human labor and long running human vision:

https://rcsnyder.github.io/open-frontier-curriculum/05-front...

https://rcsnyder.github.io/open-frontier-curriculum/05-front...

Re: GPT-6 Astra

#478

I'm sure it's going to do great on all sorts of benchmarks, but the video--the actual marketing video that if anything is incentivised to overstate things--is full of careful cuts just before it would do anything that still wouldn't actually be that impressive. It's AGI, and it's going to upload photos, or change a background slide colour. Even the people hyping it up, who believe that it's really artificial intellig…

[deleted]

Re: GPT-6 Astra

#479

What is going to become of life for those of us who do not work at AI labs and are unlikely to be hired by AI labs, despite all the years we put into learning coding, math, etc, as we were told to do? Those of us who made the mistake of studying anything other than machine learning. How will we make a living? (We don't live in a world that seems likely to distribute gains widely instead of largely to the handful of a…

[deleted]

Re: GPT-6 Astra

#480

I'm sure it's going to do great on all sorts of benchmarks, but the video--the actual marketing video that if anything is incentivised to overstate things--is full of careful cuts just before it would do anything that still wouldn't actually be that impressive. It's AGI, and it's going to upload photos, or change a background slide colour. Even the people hyping it up, who believe that it's really artificial intellig…

The benchmarks do looks good (I mean: they literally spank the latest Anthropic benchmarks of two days ago in every single benchmark) but the promotional vid is so cheesy. They decided to use the iconic Herman Miller Eames chair if I'm not mistaken: https://youtu.be/s5zyhGMMPKs And that's basically 50% of the vid looking "classy". I don't know if it's farcical but at this point --maybe I'm jaded-- I'm expecting more…

[flagged]
Post reply on HN