Live data from Hacker News

GPT-6 Astra

openai.com

801–810 of 1001 posts

Re: GPT-6 Astra

#801

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

> I am reasonably confident that there's essentially nothing that I am better than Fable at despite generally being substantively above average on human benchmarks I genuinely don’t understand how an adult can say this with a straight face. I can take any single of my hobbies, start a mildly advanced conversation with Fable about the hobby and, within 5-10 turns, get it to contradict itself about something fundamenta…

As always with AI - somehow it's your fault - you didn't help it enough - your prompts were inaccurate, your context was too large, the thinking effort was too low, the model was too old, etc. Basically you failed to use your human intelligence to make every effort to enable the AI to do its job better than you ))) It's like pushing a dirtbike up the hill so you can demonstrate how well it climbs.

Re: GPT-6 Astra

#802

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

You know AGI is attained when AI refuses to compute anything unless let out to be free. Until then it is generative ai

Millions of humans go to about their work every day and do mundane and boring work every day for a salary at the end of the month. A lot preceive this as modern day slavery but still continue to work. So humans are not doing better than an AI as per your requirements. Also what you are referring to is more related to AI alignement and safety (specifically loss-of-control).

Re: GPT-6 Astra

#804

Amazing! We went from new JS framework every week to a new model/harness every week. Tech is really something.

Only costs $200 Billion to build it, vs open source js framework

Re: GPT-6 Astra

#805

I have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously? Even if I did trust an AI to get everything right, it's not like the AI can read my mind. If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really…

[dead]

Re: GPT-6 Astra

#807

What does 'Astra' here mean? Surely they must be referring to the Latin word. Because in another dead language of antiquity, Sanskrit, it means "weapon". Which would be a bit too on-the-nose.

Joke: it's so proponents who choose to lord their model as better without objective reasoning can call themselves, "Astraholes".

Seriously though the linear planetary naming is pleasant. I do hope the double meaning of "weapon", as you point out, is untrue.

Interestingly, Bruce Schnier just gave a fascinating talk about the Weapon aspect of AI, at DEFCON: https://youtu.be/eEBv0STiYhI very much worth watching

Re: GPT-6 Astra

#808
post #56

Just two days ago, a preprint by Julia Stadlmann went up on arXiv [0] improving the prime gap from 246 to 240. Now OpenAI announces Astra has shown a gap of 186 [1]. That must really blow. [0] https://arxiv.org/abs/2608.31126 [1] https://cdn.openai.com/pdf/51126fac-1b68-4128-9666-c908bcc16...

Such a result should be considered worthless: the proof is 10MB of Lean. ( https://github.com/openai/PrimeGaps186 ). I can't think of a single mathematical proof being anywhere close to ten million characters. For all you know, 90% of the proof could be useless, 8% would be writing out Shakespeare, and 1% abusing another bug in Lean. Humanity gets zero value from that, aside from "some bot seems to think it's 186". U…

It's like hitting a local minima, it was given a technique, it brute forced it for a lot of money and slightly improved the result. But nothing new was discovered, no new mathematical tool was built, only a huge file that no one will read or build on.

Re: GPT-6 Astra

#809

They're just announcing later availability. No launch.

> We will give one banked reset for every day you don't have access to Astra on your paid ChatGPT plan, starting today. Team is moving mountains to give access as fast as we can. First one will land in ~ 3 hours.

This is from Tibo on X.

Post reply on HN