Live data from Hacker News

GPT-6 Astra

openai.com

941–950 of 1001 posts

Re: GPT-6 Astra

#941
"""OpenAI called the model a "generational leap" for areas such as cybersecurity, professional work, software engineering, and science, with the company's president Greg Brockman saying that it could eventually be seen as the arrival of artificial general intelligence"""

https://en.wikipedia.org/wiki/GPT-6_Astra

according to Sam Altman from last year, an LLM should 'solve quantum gravity', if it is to count as AGI. Did anything change since then?

https://www.businessinsider.com/sam-altman-openai-david-deut...

Re: GPT-6 Astra

#942
post #674

I think the thing I'm most excited about is the increase in _user prompting_. If I give a poorly constrained/ambiguous prompt, I don't want the model one-shotting assumptions left and right. The demos of Fable/GPT-6 are impressive, but "real AGI" should act more like a collaborator than either a peon or overachiever. It's a tough balance to get right, and although this has been possible to achieve with additional pro…

Anecdotal experiences from my external early testing of Astra: if you love Sol (like I do) and wished it was smarter at everything, but especially better at high-level tasks and discussions; I think you'll LOVE Astra. Astra retains the best parts and overall 'grounded collaborator and executor' of Sol in my testing (harness: codex CLI); while being a significant leap in capabilities & higher-level thinking. When you…

Off topic, but ooc what do you do such that you get early access to the models?

Re: GPT-6 Astra

#944
It's said Astra is based on a newly pretrained model. Personally I hope it solves the problem that GPT speaks weirdly (which can also be spotted on Claude models after Opus 4.8 but GPT's is more severe) because IMO the way GPT couldn't get how to write good code (obviously it's trined to write code which pass benchmarks, but never code which sound good, with pursuit towards simplification and code aesthetic in mind) is extremely similar with how it couldn't get how to speak like a human.

However, it seems like OpenAI didn't pay much attention on these perspectives and I didn't find if Astra could write a more elegant code, or communicate more naturally, etc., which made me somehow a little disappointed.

They indeed mentioned the code Astra delivered is closer to production grade but production-grade code is different from what I want since there can be a kind of messy code blowing up your whole architecture design with control flows nobody truly understands but just passes all tests perfectly. There is no difficulty in maintaining this kind of code because you only need to paste the problems into Codex. And we all know this sounds incorrect. I don't know if my appetite towards a good code (no matter how) is sound but I just imagined frontier labs to give more attention on this.

Note: fwiw Fable 5.1's release page says it's better at these perspectives of coding and per my experience, yes it is.

Re: GPT-6 Astra

#946

"""OpenAI called the model a "generational leap" for areas such as cybersecurity, professional work, software engineering, and science, with the company's president Greg Brockman saying that it could eventually be seen as the arrival of artificial general intelligence""" https://en.wikipedia.org/wiki/GPT-6_Astra according to Sam Altman from last year, an LLM should 'solve quantum gravity', if it is to count as AGI. D…

I think that would be more of ASI than AGI, AGI is a smart human, I'd say we're pretty close if not already there for most stuff we've been focusing on like coding. ASI is a super genius, beyond the smartest human, and we're nowhere near that. But as we've already seen with LLMs you don't need to wait for ASI to solve hard problems in math and physics.

Re: GPT-6 Astra

#947
The model's performance and efficiency seems like more evidence (bordering on the last nail in the coffin) against the assertion that inference will never be profitable to me, but I guess that's a common symptom of AI Psychosis according to the true believers in that assertion.

All my issues with its leader aside, great work OpenAI!

Re: GPT-6 Astra

#948

"""OpenAI called the model a "generational leap" for areas such as cybersecurity, professional work, software engineering, and science, with the company's president Greg Brockman saying that it could eventually be seen as the arrival of artificial general intelligence""" https://en.wikipedia.org/wiki/GPT-6_Astra according to Sam Altman from last year, an LLM should 'solve quantum gravity', if it is to count as AGI. D…

Bills are coming due, so marketing will change to fit the needs of the company's pocketbook. Please do not expect objectivity from Sam Altman or others with vested interests here. Its an impressive model, but AGI has always been a ridiculously nebulous concept used to inspire investors into giving away money.

Re: GPT-6 Astra

#949
post #681
post #641

- OpenAI claims Astra beats all benchmarks (compared to Fable and Opus, except "Humanity's Last Exam (w/ tools)"): https://openai.com/index/gpt-6-astra/ - Artificial Analysis scores Astra (max effort) as 61 points on intelligence, behind Opus 5. https://artificialanalysis.ai/models/gpt-6-astra Who is wrong here? Some benchmark results in Astra page for Fable and Opus are blank (-). What is Artificial Analysis intelli…

I really, really don't find the Artificial Analysis Intelligence Index credible anymore. It's some weighted score of benchmarks, and benchmarks increasingly don't reflect how good a model is. That should be obvious if you compare Gemini 3.8 Flash (which is an _excellent_ model especially for its price and TPS!! but 10min of prompting in any harness) will tell you it's nowhere near close to Sol/Astra. But AA scores Ge…

> but 10min of prompting in any harness) will tell you it's nowhere near close to Sol/Astra.

I code in both every day a lot and it is not obvious to me 3.8 is far behind

Re: GPT-6 Astra

#950
I'm sure this will be a great model.

Personally, I'm far away from screaming 'AGI is here!' from the rooftops, until jaggedness and silly mistakes disappear at the very least . (what is going on with that Mario Kart game...) So many benchmarks are 'best of x tries' or using very specific harnesses. AGI would not need a babysitter.

Honestly, even being able to do simple tasks like summarization or basic knowledge work without the constant paranoia of unforseen failure would be remarkable and useful.

Post reply on HN