https://en.wikipedia.org/wiki/GPT-6_Astra
according to Sam Altman from last year, an LLM should 'solve quantum gravity', if it is to count as AGI. Did anything change since then?
https://www.businessinsider.com/sam-altman-openai-david-deut...
941–950 of 1001 posts
https://en.wikipedia.org/wiki/GPT-6_Astra
according to Sam Altman from last year, an LLM should 'solve quantum gravity', if it is to count as AGI. Did anything change since then?
https://www.businessinsider.com/sam-altman-openai-david-deut...
I think the thing I'm most excited about is the increase in _user prompting_. If I give a poorly constrained/ambiguous prompt, I don't want the model one-shotting assumptions left and right. The demos of Fable/GPT-6 are impressive, but "real AGI" should act more like a collaborator than either a peon or overachiever. It's a tough balance to get right, and although this has been possible to achieve with additional pro…
Anecdotal experiences from my external early testing of Astra: if you love Sol (like I do) and wished it was smarter at everything, but especially better at high-level tasks and discussions; I think you'll LOVE Astra. Astra retains the best parts and overall 'grounded collaborator and executor' of Sol in my testing (harness: codex CLI); while being a significant leap in capabilities & higher-level thinking. When you…
It's time they stop picking the low hanging fruit and start folding my laundry!
However, it seems like OpenAI didn't pay much attention on these perspectives and I didn't find if Astra could write a more elegant code, or communicate more naturally, etc., which made me somehow a little disappointed.
They indeed mentioned the code Astra delivered is closer to production grade but production-grade code is different from what I want since there can be a kind of messy code blowing up your whole architecture design with control flows nobody truly understands but just passes all tests perfectly. There is no difficulty in maintaining this kind of code because you only need to paste the problems into Codex. And we all know this sounds incorrect. I don't know if my appetite towards a good code (no matter how) is sound but I just imagined frontier labs to give more attention on this.
Note: fwiw Fable 5.1's release page says it's better at these perspectives of coding and per my experience, yes it is.
"""OpenAI called the model a "generational leap" for areas such as cybersecurity, professional work, software engineering, and science, with the company's president Greg Brockman saying that it could eventually be seen as the arrival of artificial general intelligence""" https://en.wikipedia.org/wiki/GPT-6_Astra according to Sam Altman from last year, an LLM should 'solve quantum gravity', if it is to count as AGI. D…
All my issues with its leader aside, great work OpenAI!
"""OpenAI called the model a "generational leap" for areas such as cybersecurity, professional work, software engineering, and science, with the company's president Greg Brockman saying that it could eventually be seen as the arrival of artificial general intelligence""" https://en.wikipedia.org/wiki/GPT-6_Astra according to Sam Altman from last year, an LLM should 'solve quantum gravity', if it is to count as AGI. D…
- OpenAI claims Astra beats all benchmarks (compared to Fable and Opus, except "Humanity's Last Exam (w/ tools)"): https://openai.com/index/gpt-6-astra/ - Artificial Analysis scores Astra (max effort) as 61 points on intelligence, behind Opus 5. https://artificialanalysis.ai/models/gpt-6-astra Who is wrong here? Some benchmark results in Astra page for Fable and Opus are blank (-). What is Artificial Analysis intelli…
I really, really don't find the Artificial Analysis Intelligence Index credible anymore. It's some weighted score of benchmarks, and benchmarks increasingly don't reflect how good a model is. That should be obvious if you compare Gemini 3.8 Flash (which is an _excellent_ model especially for its price and TPS!! but 10min of prompting in any harness) will tell you it's nowhere near close to Sol/Astra. But AA scores Ge…
I code in both every day a lot and it is not obvious to me 3.8 is far behind
Personally, I'm far away from screaming 'AGI is here!' from the rooftops, until jaggedness and silly mistakes disappear at the very least . (what is going on with that Mario Kart game...) So many benchmarks are 'best of x tries' or using very specific harnesses. AGI would not need a babysitter.
Honestly, even being able to do simple tasks like summarization or basic knowledge work without the constant paranoia of unforseen failure would be remarkable and useful.