Live data from Hacker News

GPT-6 Astra

openai.com

501–510 of 1001 posts

Re: GPT-6 Astra

#501

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

It doesn’t even know what day it is unless it’s told. Statelessness is never going to be ”general intelligence” in my book, and the concept of ”memory” in models are laughably bad today. Then again, who cares, AGI means nothing anymore, it’s a term for marketing only and has no technical or scientific meaning.

Re: GPT-6 Astra

#503
post #58

Just two days ago, a preprint by Julia Stadlmann went up on arXiv [0] improving the prime gap from 246 to 240. Now OpenAI announces Astra has shown a gap of 186 [1]. That must really blow. [0] https://arxiv.org/abs/2608.31126 [1] https://cdn.openai.com/pdf/51126fac-1b68-4128-9666-c908bcc16...

Such a result should be considered worthless: the proof is 10MB of Lean. ( https://github.com/openai/PrimeGaps186 ). I can't think of a single mathematical proof being anywhere close to ten million characters. For all you know, 90% of the proof could be useless, 8% would be writing out Shakespeare, and 1% abusing another bug in Lean. Humanity gets zero value from that, aside from "some bot seems to think it's 186". U…

The proof of the classification of finite simple groups is bigger than that.

Re: GPT-6 Astra

#504

Lol their page finally loaded. They added an example scenario of "Filling in Form 1040" - which made me laugh out loud. That is indeed something most US citizens cannot accurately do even with expensive proprietary tax software services. Kind of a Hitchhiker's Guide to the Galaxy meme but where the tax code is so complicated we're implementing powerful AIs to be able to do it (hopefully) right.

i tried to get claude to do my taxes for last year and it refused :(

now that i'm a gpt subscriber maybe I'll have luck when i'm filing next year

Re: GPT-6 Astra

#506

I'm sure it's going to do great on all sorts of benchmarks, but the video--the actual marketing video that if anything is incentivised to overstate things--is full of careful cuts just before it would do anything that still wouldn't actually be that impressive. It's AGI, and it's going to upload photos, or change a background slide colour. Even the people hyping it up, who believe that it's really artificial intellig…

The tone of the marketing video is a bit irritating to me as someone who has been laid off and feels cheated and fearful of AI. It shows people who seem to have very full and rich lives, and the reason they do is because they use ChatGPT. These are the people smart enough to say things like "do what needs to be done", or "change the background to make it look better"--insights like these are why they make the big buc…

The world is changing. Wont help if you keep stuck in the old world.

Re: GPT-6 Astra

#507
post #317

Is anyone else just exhausted by the pace of all this. The models change constantly and relentlessly and so does the pricing, basically weekly at this point between all the labs. It feels nearly impossible to have any rigorous approach when choosing a particular model and price point for a task and more like blindly picking one. The time period needed to actually get familiar with various models to a degree you can i…

Hey, I'm on the team at LiteLLM that's building the auto-router and our goal right now is to abstract that decision making away from the end user. The biggest thing we're trying to figure out right now is how do we do that without frustrating the end user - as a developer myself I would hate for my agent to be dumbed down below the threshold needed to complete a task.

In theory though, there is a minimum viable model for any given task, and we think that is a problem that the big labs will avoid because they profit from charging more per task. We're trying heuristic and LLM-based approaches but it's still a work in progress, so if this is something you'd be interested in trying would highly recommend trying ours out -- any and all feedback at this point is extremely valuable to us.

https://docs.litellm.ai/docs/proxy/auto_routing

Re: GPT-6 Astra

#508

I'm sure it's going to do great on all sorts of benchmarks, but the video--the actual marketing video that if anything is incentivised to overstate things--is full of careful cuts just before it would do anything that still wouldn't actually be that impressive. It's AGI, and it's going to upload photos, or change a background slide colour. Even the people hyping it up, who believe that it's really artificial intellig…

The tone of the marketing video is a bit irritating to me as someone who has been laid off and feels cheated and fearful of AI. It shows people who seem to have very full and rich lives, and the reason they do is because they use ChatGPT. These are the people smart enough to say things like "do what needs to be done", or "change the background to make it look better"--insights like these are why they make the big buc…

Because they earned it with their strong entrepreneurial spirit and grit.

Haven’t you learned anything?

Re: GPT-6 Astra

#509

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

> even with Fable, I've been at the point personally where I am reasonably confident that there's essentially nothing that I am better than Fable at despite generally being substantively above average on human benchmarks

If I asked you to write fiction, you'd be much better at keeping track of which characters knew which facts.

Re: GPT-6 Astra

#510
post #247

I want to take a step back: So, this is GPT-6 -- the natural number version release comparable to GPT-4 and GPT-5 from the past few years. The ARC-AGI-3 score is obviously impressive at 99.9% (we'll need to wait for more details on how they used the response API harness on GPT-6 Astra, wrt reasoning retention and compaction), but every other benchmarks seems to be a relatively modest improvement, comparable with any…

• 98.6% on ARC-AGI-3

• 97.6% on frontier math

• 95.9% on CAD

• 100% on ExploitBench

Nothing modest about it

Post reply on HN