Live data from Hacker News

GPT-6 Astra

openai.com

311–320 of 1001 posts

Re: GPT-6 Astra

#311
post #230

I want to take a step back: So, this is GPT-6 -- the natural number version release comparable to GPT-4 and GPT-5 from the past few years. The ARC-AGI-3 score is obviously impressive at 99.9% (we'll need to wait for more details on how they used the response API harness on GPT-6 Astra, wrt reasoning retention and compaction), but every other benchmarks seems to be a relatively modest improvement, comparable with any…

It's like my RPG character putting every points to one single trait. I'll one shot everything alive but will instantly die if accidentally drink water with 6.9 pH.

Re: GPT-6 Astra

#312

The FrontierCode 1.1 Extended benchmark is the only benchmark that aligns with my actual LLM experiences and Astra isn't significantly better or cheaper. All this celebration, and yet it's only on-par with an already existing model? I don't get it.

If your benchmark shows Opus 5 winning, I really question the validity of it.

Re: GPT-6 Astra

#313
post #78

ARC AGI-3 saturated by Astra! https://arcprize.org/leaderboard

I think it could indicate that "semi-private" dataset likely leaked to their training data.

A dataset being as popular as their's is will contaminate the data just by people discussing it and creating their own public test sets of similar problems.

Still, probably not that much compared to employees targeting it.

Re: GPT-6 Astra

#314
post #292

Is anyone else just exhausted by the pace of all this. The models change constantly and relentlessly and so does the pricing, basically weekly at this point between all the labs. It feels nearly impossible to have any rigorous approach when choosing a particular model and price point for a task and more like blindly picking one. The time period needed to actually get familiar with various models to a degree you can i…

That's how cutting edge tech has always worked.

Imagine buying a shiny new PC in the 90s only to see it become practically obsolete within a year.

Re: GPT-6 Astra

#315
post #267

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

Take it from the mouth of the creator of ARC-AGI: When we released ARC 3, I got asked, "when do you think a frontier model will saturate it?", and I answered "in about a year, though it depends on how much it gets explicitly targeted" That was 6 months ago, so the progress that Astra represents happened about 2x faster than I anticipated. I think the speed of progress will surprise a lot of people, and what the new m…

>When we released ARC 3, I got asked, "when do you think a frontier model will saturate it?", and I answered "in about a year, though it depends on how much it gets explicitly targeted"

You're treating an off-hand comment by an ARC 3 researcher as some sort of a precise AI capability acceleration benchmark. Can we leave casual anecdotes (even from researchers) out of the discussions please?

Re: GPT-6 Astra

#317
post #267

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

Take it from the mouth of the creator of ARC-AGI: When we released ARC 3, I got asked, "when do you think a frontier model will saturate it?", and I answered "in about a year, though it depends on how much it gets explicitly targeted" That was 6 months ago, so the progress that Astra represents happened about 2x faster than I anticipated. I think the speed of progress will surprise a lot of people, and what the new m…

I feel like AGI's definition got watered down, and these tests do not cover the original definition, what is your definition and thoughts on aligning with what all of us understood from the original claim?

I feel like this test is just helping someone like Sam Altman pretend like he implemented AGI as originally pitched for an IPO when in fact, he has not. Shameful.

> AGI is essentially the equivalent of a median human that could be hired as a remote co-worker... capable of performing any task that one would be satisfied with a remote colleague doing via a computer.

- Sam Altman on AGI

Re: GPT-6 Astra

#318
Ok, but can I bring GPT-6 in as an agent as a software engineer, tell it to talk to these people and have it start solving engineering problems and continue on for a full year career wise?

maybe call it EngEmployeeBench

Re: GPT-6 Astra

#319

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

It has to pass the Turing test

I'm barely holding it together here so you don't get the full spiel, but a quick skim of Turing's paper clarifies that it was never about a binary test. https://courses.cs.umbc.edu/471/papers/turing.pdf Specifically sections 1 & 6 dispell the common myths, and the conclusion is also quite powerful.

Smart guy, that Turing. I wish he were still around... Linus but 114 years old and with 8 of that as the chair of a federated EU, kept alive by his own positive impact on dissolving the cold war into even more of a scientific boom. Would crazy helpful as we try to navigate the interesting times within which we have been damned.

A comforting thought, almost?

Post reply on HN