ARC AGI-3 saturated by Astra! https://arcprize.org/leaderboard
The no-reasoning version scores 35% while the low reasoning one scores 17%? What?
GPT-6 Astra
711–720 of 1001 posts
Re: GPT-6 Astra
#712I think the thing I'm most excited about is the increase in _user prompting_. If I give a poorly constrained/ambiguous prompt, I don't want the model one-shotting assumptions left and right. The demos of Fable/GPT-6 are impressive, but "real AGI" should act more like a collaborator than either a peon or overachiever. It's a tough balance to get right, and although this has been possible to achieve with additional pro…
I don't really agree. The thing that makes Fable feel like an actual collaborator is its ability to sus out your real intent when you give ambiguous instructions. It's really good at it.
I watched some reviews today and came way with the impression that Astra is not better than Sol in this regard. You still have to be very specific with your instructions. For example, you can say "why is it not committed yet?" and it will give you an explanation and say it's actually ready to be committed. But it won't commit unless you explicitly say so.
That sounds like a very tedious way of working with AI agents, but I understand some people want a high level of control.
Re: GPT-6 Astra
#713The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…
Most things in the real world probably. I’m not saying AI can’t do it, but currently it’s bad. Try send an image of the inside of a broken toaster and how to fix. It’s laughable. Again, not saying AI will never do it, but am saying there are definitely large holes in knowledge.
Re: GPT-6 Astra
#714The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…
Re: GPT-6 Astra
#715Oops, shots fired. A direct attack on the vibe coded app market. Replit, Lovable, etc.
Re: GPT-6 Astra
#716I am really confused on how it can saturate ARC-AGI but still perform poorly on aggregated benchmarks: https://artificialanalysis.ai/models Perhaps if it was allowed this custom harness for all benchmarks it would similarily saturate?
Re: GPT-6 Astra
#717Re: GPT-6 Astra
#718Re: GPT-6 Astra
#719I want to take a step back: So, this is GPT-6 -- the natural number version release comparable to GPT-4 and GPT-5 from the past few years. The ARC-AGI-3 score is obviously impressive at 99.9% (we'll need to wait for more details on how they used the response API harness on GPT-6 Astra, wrt reasoning retention and compaction), but every other benchmarks seems to be a relatively modest improvement, comparable with any…