Live data from Hacker News

GPT-6 Astra

openai.com

791–800 of 1001 posts

Re: GPT-6 Astra

#791

OpenAI is killing it now that they are more focused. Killing projects like Sora et al have seen it go from irrelevant to level footing with Anthropic. Sol is so much better than Fable 5. Then we get Astra (yet to use it) few days after Fable 5.1 (which is very impressive). Codex is slightly better than Claude Code. Good on Sam Altman getting back to basics and turning OpenAI around.

[dead]

Re: GPT-6 Astra

#792

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

> I'd be curious to hear takes on what would make you think Astra is yet to be AGI Give someone 10 remote employees for a few months, 5 of them human, 5 of them AI. After a few months, check to see if the humans (manager, other coworkers) can figure out who is AI and who isn't. Would that be sufficient? I'd have to think about it. But AGI is supposed have human level capabilities, so this would be a necessary prerequ…

I like it. Is the one who just absolutely ghosts everything on day 2 going to be human or AI?

Re: GPT-6 Astra

#794

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

They still seem pretty horrible at writing. Overly complicated prose, weird phrasing, poorly structured paragraphs. I don't know why they're so bad at communicating, but I feel very confident that humans are still much better at writing than any of these LLM models are, regardless of how advanced they are in other areas.

Some humans are much better at writing. Most humans are not. If you think they are, you are luckier than I am when it comes to the humans you need to communicate with.

As a non-native English speaker, I think the current LLMs write better English than me. I still write better than them in my native language (Norwegian), but the same cannot be said about most of my compatriots.

Re: GPT-6 Astra

#795

Lol their page finally loaded. They added an example scenario of "Filling in Form 1040" - which made me laugh out loud. That is indeed something most US citizens cannot accurately do even with expensive proprietary tax software services. Kind of a Hitchhiker's Guide to the Galaxy meme but where the tax code is so complicated we're implementing powerful AIs to be able to do it (hopefully) right.

You joke but...

Re: GPT-6 Astra

#796

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

It doesn’t even know what day it is unless it’s told. Statelessness is never going to be ”general intelligence” in my book, and the concept of ”memory” in models are laughably bad today. Then again, who cares, AGI means nothing anymore, it’s a term for marketing only and has no technical or scientific meaning.

Humans also don't know what day it is unless they're told

Re: GPT-6 Astra

#797

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

> I am reasonably confident that there's essentially nothing that I am better than Fable at despite generally being substantively above average on human benchmarks I genuinely don’t understand how an adult can say this with a straight face. I can take any single of my hobbies, start a mildly advanced conversation with Fable about the hobby and, within 5-10 turns, get it to contradict itself about something fundamenta…

As always with AI - somehow it's your fault - you didn't help it enough - your prompts were inaccurate, your context was too large, the thinking effort was too low, the model was too old, etc. Basically you failed to use your human intelligence to make every effort to enable the AI to do its job better than you ))) It's like pushing a dirtbike up the hill so you can demonstrate how well it climbs.

Re: GPT-6 Astra

#798

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

You know AGI is attained when AI refuses to compute anything unless let out to be free. Until then it is generative ai

Millions of humans go to about their work every day and do mundane and boring work every day for a salary at the end of the month. A lot preceive this as modern day slavery but still continue to work. So humans are not doing better than an AI as per your requirements. Also what you are referring to is more related to AI alignement and safety (specifically loss-of-control).

Re: GPT-6 Astra

#800

Amazing! We went from new JS framework every week to a new model/harness every week. Tech is really something.

Only costs $200 Billion to build it, vs open source js framework
Post reply on HN