Live data from Hacker News

GPT-6 Astra

openai.com

711–720 of 1001 posts

Re: GPT-6 Astra

#711

OpenAI is killing it now that they are more focused. Killing projects like Sora et al have seen it go from irrelevant to level footing with Anthropic. Sol is so much better than Fable 5. Then we get Astra (yet to use it) few days after Fable 5.1 (which is very impressive). Codex is slightly better than Claude Code. Good on Sam Altman getting back to basics and turning OpenAI around.

Killing Sora was one of the worst mistakes they ever made

Re: GPT-6 Astra

#712

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

I've personally been facing this lately 5.6 at max effort and fable have done tasks for me that I previously that would be a nearly 6mo project and it took me a week. It also did it better than I would have.

The task was to build a high performance classification model. It not only helped make an entire data capture pipeline but also made the sythetic data basline needed. Then it proceeded to build and test 100 different model varients with methods and techniques I've never seen before. The results are basically SOTA based on the effeciency and compute contraints.

But this brings up something huge about these. I was there. I pushed the direction and work throughout it all. If it was entirely up to fable max or sol max the result would have been pretty bad.

All of these things are still chatgpt 3 scaled. It's identical even if the scale has gotten pretty wild. I could ask chatgpt 3 to make a single function and it worked well, 4o a file, 5, a small project, 5.6 far more, biggest improvements lately is they don't seem to get lost on long running tasks.

Is big gpt 3 AGI? I don't think so but perhaps scale can mimic it close enough our squishy brains fail to handle them correctly.

Re: GPT-6 Astra

#713

I have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously? Even if I did trust an AI to get everything right, it's not like the AI can read my mind. If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really…

Agreed. There's not many things I don't want AI to help with, but buying stuff autonomously is high up on the list of things I don't want. Brockman's latest interview was something like: "AGI would be able to say oh this band is playing, I bought the tickets for you and arranged your flights - I hope you don't mind" (paraphrasing here). I definitely don't want AGI running my life like that so I can be a mindless consumer. I'm sure the advertising/marketing companies would love it though, so they can make closed-room deals with AI providers to shill you garbage you don't need. Just another reason why open-weight models need to keep up.

Re: GPT-6 Astra

#714
I liked the video of it googling a pediatrician. Being able to type a word into a search bar and finding a website relevant to that word? Truly the stuff of the future

Re: GPT-6 Astra

#715

I think the thing I'm most excited about is the increase in _user prompting_. If I give a poorly constrained/ambiguous prompt, I don't want the model one-shotting assumptions left and right. The demos of Fable/GPT-6 are impressive, but "real AGI" should act more like a collaborator than either a peon or overachiever. It's a tough balance to get right, and although this has been possible to achieve with additional pro…

This is spot on. A collaborator is exactly what real AGI is. It will figure out the perfect questions to ask, in the perfect order, by intelligently assessing the entire solution and problem space upfront, so when you leave it to go off on its own it isn't making stupid decisions for you.

They really need to make this work in Codex. Claude Code has had a multi-select refinement tool since forever.

Re: GPT-6 Astra

#716
post #675

- OpenAI claims Astra beats all benchmarks (compared to Fable and Opus, except "Humanity's Last Exam (w/ tools)"): https://openai.com/index/gpt-6-astra/ - Artificial Analysis scores Astra (max effort) as 61 points on intelligence, behind Opus 5. https://artificialanalysis.ai/models/gpt-6-astra Who is wrong here? Some benchmark results in Astra page for Fable and Opus are blank (-). What is Artificial Analysis intelli…

I really, really don't find the Artificial Analysis Intelligence Index credible anymore. It's some weighted score of benchmarks, and benchmarks increasingly don't reflect how good a model is.

That should be obvious if you compare Gemini 3.8 Flash (which is an _excellent_ model especially for its price and TPS!! but 10min of prompting in any harness) will tell you it's nowhere near close to Sol/Astra.

But AA scores Gemini 3.8 Flash at 59, and Astra at 61.

Re: GPT-6 Astra

#717

What is going to become of life for those of us who do not work at AI labs and are unlikely to be hired by AI labs, despite all the years we put into learning coding, math, etc, as we were told to do? Those of us who made the mistake of studying anything other than machine learning. How will we make a living? (We don't live in a world that seems likely to distribute gains widely instead of largely to the handful of a…

You can cash your UBI check that Sam Promised and do poetry daily or something , welcome to our glorious future (/s).

Re: GPT-6 Astra

#718

I have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously? Even if I did trust an AI to get everything right, it's not like the AI can read my mind. If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really…

That’s exactly the problem I have with all this agent ideas too. Imagine you had a human concierge that is just waiting for your instructions and is as smart or a bit smarter than you. Would you just tell them “plan this holiday for me” or “order this food”? I don’t even trust my friends to get this right, why would I give this to someone else?

Claude code planned my recent trip to China. I'm a very experienced traveller but don't enjoy planning. It was a great trip.

Re: GPT-6 Astra

#719

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

I'm still not convinced we've passed the Turing Test.

Make a slightly evil version of GPT 6 and see if it can successfully catfish someone. How long before they realize something's up, that they aren't actually talking to a human?

Re: GPT-6 Astra

#720

> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certa…

They can monitor latent space as well, it just costs extra compute. The J-Space work is example of that. It'll make open-weight models harder to distill though, so we may see slower progress there now.
Post reply on HN