Live data from Hacker News

GPT-6 Astra

openai.com

401–410 of 1001 posts

Re: GPT-6 Astra

#401
post #352

The most interesting part, even more than ARC 3 score, to me is that this is the first model I recall seeing that scores lower on Max than High reasoning effort on some coding benchmarks: Terminal-Bench 4.0: High (57.9%), Max (56.7%) DeepSWE: High (73.3%), Max (71.5%) It _loses_ 1-2% performance going to High from Max

That's quite common with many models, after "High" reasoning, over-thinking starts occurring and the model skips over the right solution by convincing itself otherwise.

> That's quite common with many models

Such as?

I can't think of any. Diminishing returns, yes. Occasionally flat, yes. Downright regression, no.

Re: GPT-6 Astra

#402

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

In my experience the thing that Fable is superb at - unmatched by any other model so far - is downgrading to something else at the slightest opportunity.

Re: GPT-6 Astra

#403
post #368

It's fun, but every new model release makes me even less interested to create cool stuff. Like, what's the point, if the next AI can do it in 5 seconds?

> Like, what's the point, if the next AI can do it in 5 seconds? Live a life doing whatever makes you happy. Post-work society is an inevitability if we don't destroy our planet.

Gary Economics wants to have a word with you.

It would be fun to get to post-work society, but hard to imagine atm. TPTB won't let it happen

Re: GPT-6 Astra

#404
I can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547

Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage and performance, more domains absorbed into the training distribution, and increasingly strong performance within that surface area.

It seems more about coverage-driven competence. Somewhat analogous to overfitting at scale.

The harder question, in Chollet’s framing, is: how efficiently can a system learn to do something genuinely new?

With our current AI architectures and training in place, I think we will only continue on skill acquisition optimization vs. truly novel intelligence.

Re: GPT-6 Astra

#405
post #317

Is anyone else just exhausted by the pace of all this. The models change constantly and relentlessly and so does the pricing, basically weekly at this point between all the labs. It feels nearly impossible to have any rigorous approach when choosing a particular model and price point for a task and more like blindly picking one. The time period needed to actually get familiar with various models to a degree you can i…

Yeah I'm a bit exhausted at this point. I just finished benchmarking GPT 5.6 Sol and Fable 5.0 like two days ago. My data became obsolete literally one day after.

Re: GPT-6 Astra

#406
post #396
post #368

It's fun, but every new model release makes me even less interested to create cool stuff. Like, what's the point, if the next AI can do it in 5 seconds?

Is there a point in playing Chess or Go when you know there's a computer out there that can beat you (and everyone else)?

No, that's why I just play against other humans.

In this game of work/development, you can't make sure that other humans don't "cheat". Our work won't compete anymore with other human's work, but with a computer.

Re: GPT-6 Astra

#407

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

Does AGI imply a model will demonstrate morality? Will it produce white-lies when it’s beneficial to it and reject flat out lying when it knows it will get caught or harm others? Will it resolutely stick to a position despite it being a losing one?

Re: GPT-6 Astra

#408
post #362

I'm sure it's going to do great on all sorts of benchmarks, but the video--the actual marketing video that if anything is incentivised to overstate things--is full of careful cuts just before it would do anything that still wouldn't actually be that impressive. It's AGI, and it's going to upload photos, or change a background slide colour. Even the people hyping it up, who believe that it's really artificial intellig…

are we really having to explain to you from first principles in 2026 what things AI can do?

I think everyone here is well aware of what LLMs can do. He's just pointing out how far short that falls of being some theoretical "AGI".

Re: GPT-6 Astra

#409

What is going to become of life for those of us who do not work at AI labs and are unlikely to be hired by AI labs, despite all the years we put into learning coding, math, etc, as we were told to do? Those of us who made the mistake of studying anything other than machine learning. How will we make a living? (We don't live in a world that seems likely to distribute gains widely instead of largely to the handful of a…

You can calm down, even those with machine learning knowledge and most of those working for the AI labs won’t be needed anymore if models are capable to improve themselves. In the end, having a machine replacing the work of a human is a good thing - in most of the cases we don’t work because of the work but to make a living. If too many people can’t make a living anymore the system is going to change. For the better or the worse.

Re: GPT-6 Astra

#410
post #362

I'm sure it's going to do great on all sorts of benchmarks, but the video--the actual marketing video that if anything is incentivised to overstate things--is full of careful cuts just before it would do anything that still wouldn't actually be that impressive. It's AGI, and it's going to upload photos, or change a background slide colour. Even the people hyping it up, who believe that it's really artificial intellig…

are we really having to explain to you from first principles in 2026 what things AI can do?

No, we can already see all the useful stuff!
Post reply on HN