Live data from Hacker News

GPT-6 Astra

openai.com

251–260 of 1001 posts

Re: GPT-6 Astra

#251
post #78

ARC AGI-3 saturated by Astra! https://arcprize.org/leaderboard

ARC has their own writeup on the result, which offers some nuance. https://arcprize.org/blog/astra tl;dr it's 62% when apples-to-apples to other models, which is still notable.

ARC's harness is just straight up broken. No serious harness removes reasoning context between each step. Not only does this significantly lower performance over all reasoning LLMs, but it also increase cost as you destroy the cache on every turn. Tossing the oldest entry when context fills up instead of using compaction is equally bad with the same issues.

Re: GPT-6 Astra

#252
Maybe it is AGI and they didn't benchmax it or it is not and is worse then 5.6 sol, which if true would just be sad

Re: GPT-6 Astra

#253

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

It's "harnessmaxxing" all the way down. AI benchmark scene is exhibit A for Goodhart's law.

Re: GPT-6 Astra

#254
Anthropic should prep 5.2 and 5.3 at the same time, release 5.2, wait for Google to release their shit in a day or two later than then release 5.3 just to fuck with them :)

Re: GPT-6 Astra

#255
post #94

> GPT‑6 Astra is rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and AWS. Not on Azure? If so, that's a big deal.

It's on Azure also, here is their announcement: https://azure.microsoft.com/en-us/blog/gpt-6-astra-frontier-...

Although I was also surprised they didn't have some type of contractual obligation to list that alongside AWS.

Re: GPT-6 Astra

#256
I'm sure it's going to do great on all sorts of benchmarks, but the video--the actual marketing video that if anything is incentivised to overstate things--is full of careful cuts just before it would do anything that still wouldn't actually be that impressive.

It's AGI, and it's going to upload photos, or change a background slide colour. Even the people hyping it up, who believe that it's really artificial intelligence in every sense of the word, couldn't get it to do more than that.

This is farcical.

Re: GPT-6 Astra

#257

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

It has to pass the Turing test

Re: GPT-6 Astra

#258
The ARC-AGI-3 score is an incredible feat. It needed to effectively create a symbolic world model from scratch to solve the games.

If you've played the games firsthand, you know what an accomplishment this is. The "games" feel like a weird conduit to a lower level of your brain, where you move pieces to a specific place because it just "feels" right. For AI to nail it better than a human speaks to some magic happening underneath.

Looking forward to ARC-AGI-4,5,6 and slowly chipping away at the remaining problem sets.

Re: GPT-6 Astra

#259
post #227

I want to take a step back: So, this is GPT-6 -- the natural number version release comparable to GPT-4 and GPT-5 from the past few years. The ARC-AGI-3 score is obviously impressive at 99.9% (we'll need to wait for more details on how they used the response API harness on GPT-6 Astra, wrt reasoning retention and compaction), but every other benchmarks seems to be a relatively modest improvement, comparable with any…

> If this is truly AGI (subject to one's definition of AGI still) Scoring well in a benchmark that's called AGI does not make an LLM AGI.

Hey now! Keep your reason out of their marketin^H^H lies!

Re: GPT-6 Astra

#260
Benchmark wise 5% improvement over Sol in coding tasks and a 2-3% improvement over Fable 5.1 seems pretty disappointing, but maybe it is actually much better in real world usage. Let’s see
Post reply on HN