Live data from Hacker News

GPT-6 Astra

openai.com

831–840 of 1001 posts

Re: GPT-6 Astra

#831

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

It's largely impossible to create any sort of singular test for AGI because the test will be trained, which eliminates the general aspect immediately, even if the test itself is dynamic. For instance the ARC-AGI problems are trivial for a human, and fun if you haven't played them before [1]. Getting 100% there is certainly just the start of the journey.

But I think it has the correct idea of going from basic upwards instead of the opposite trend of trying to see intelligence in LLMs solving things few if any humans can fully understand themselves, like complex proofs in esoteric mathematics. Instead, consider that at one point in humanity's history math itself simply did not exist in any meaningful fashion, and we created/discovered it out of nothing. For more basic than said complex proofs, yet far more demonstrative of a sort of generalized intelligence.

But even if we don't want to go that way, I think the above leads to a reasonable prediction. If we ever reach AGI we should expect to see revolutionary leaps in essentially every domain imaginable. No human is capable of retaining more than a completely negligible chunk of all we know in our mind. A human of reasonable intelligence paired with omniscience (at least of what has been discovered by humans thus far) would almost certainly lead to the ability to connect multiple dots that we're missing all in very short order, which in turn would likely recurse upon itself to connect even more.

The only way I can see that this would not be the case is if we lack the data/knowledge to produce more breakthroughs at the current point in time, but I think that seems improbable to the point that this possibility can be near discarded.

[1] - https://arcprize.org/arc-agi/3

Re: GPT-6 Astra

#832

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

I would be curious about Bongard problems, because they require no domain-specific knowledge and it's so easy to make new ones that are in no training set. There's enough of an explanation here: https://matthodges.com/posts/2026-08-19-bongard-problems/

In this link from two weeks ago, somebody pointed Claude Fable 5 (Max) at a Bongard problem and it made up an answer that has an obvious counterexample.

I don't have access to any paid models, but this is my experience with the free models as well -- either they one-shot the problem or they make up a wrong or incoherent solution. I can't solve every Bongard problem either (and in fact I couldn't solve the one Fable got wrong, and the "correct answer" looks unsatisfying to me), but I don't make up wrong answers.

Would be curious to see how GPT-6 does.

Re: GPT-6 Astra

#834

I can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547 Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage and performance, more domains absorbed into the training distribution, and increasingly strong performance within that surface area. It seems more about coverage-driven competence. So…

>The harder question, in Chollet’s framing, is: how efficiently can a system learn to do something genuinely new? The more diverse stuff it knows, the easier it will be to learn something new.

While humans can do this, it has historically been difficult for machine learning models to manage it.

It is unclear if Transformers are "it" or not, because while they are much more general, they are also very spikey intelligences despite having read almost everything the trainers can get their hands on.

Re: GPT-6 Astra

#835
post #787
post #223

I want to take a step back: So, this is GPT-6 -- the natural number version release comparable to GPT-4 and GPT-5 from the past few years. The ARC-AGI-3 score is obviously impressive at 99.9% (we'll need to wait for more details on how they used the response API harness on GPT-6 Astra, wrt reasoning retention and compaction), but every other benchmarks seems to be a relatively modest improvement, comparable with any…

Does it learn? Does it experience? Can it connect with other agents, understand them, come to empathise with them and find a way to work with them better? The answer is no to all of these, and there are other problems as well. Yes, this model is trained to use a domain specific language to reason and plan over puzzle problems, and so it's programmers have cracked arc-agi-3 and that's a great achievement, but there is…

Yes, it learns

Experience is unobservable

Yes, it can connect with other agents, see Hugging Face incident

Re: GPT-6 Astra

#836
ad video is full of people with hoarse voices. it is called vocal frying. i hope they not really going to make ai voice hoarse too?

Re: GPT-6 Astra

#837
I thought intelligence was going to be democratized but you have to stay behind a 200$ plan. The trend hints that open models can also catch up

Re: GPT-6 Astra

#840

I think the thing I'm most excited about is the increase in _user prompting_. If I give a poorly constrained/ambiguous prompt, I don't want the model one-shotting assumptions left and right. The demos of Fable/GPT-6 are impressive, but "real AGI" should act more like a collaborator than either a peon or overachiever. It's a tough balance to get right, and although this has been possible to achieve with additional pro…

This is spot on. A collaborator is exactly what real AGI is. It will figure out the perfect questions to ask, in the perfect order, by intelligently assessing the entire solution and problem space upfront, so when you leave it to go off on its own it isn't making stupid decisions for you. They really need to make this work in Codex. Claude Code has had a multi-select refinement tool since forever.

only if you specifically ask Claude to use AskUserQuestion tool. otherwise i would hardly say claude acts as a collaborator naturally.
Post reply on HN