Live data from Hacker News

GPT-6 Astra

openai.com

941–950 of 1001 posts

Re: GPT-6 Astra

#941

"""OpenAI called the model a "generational leap" for areas such as cybersecurity, professional work, software engineering, and science, with the company's president Greg Brockman saying that it could eventually be seen as the arrival of artificial general intelligence""" https://en.wikipedia.org/wiki/GPT-6_Astra according to Sam Altman from last year, an LLM should 'solve quantum gravity', if it is to count as AGI. D…

I think that would be more of ASI than AGI, AGI is a smart human, I'd say we're pretty close if not already there for most stuff we've been focusing on like coding. ASI is a super genius, beyond the smartest human, and we're nowhere near that. But as we've already seen with LLMs you don't need to wait for ASI to solve hard problems in math and physics.

Re: GPT-6 Astra

#942
The model's performance and efficiency seems like more evidence (bordering on the last nail in the coffin) against the assertion that inference will never be profitable to me, but I guess that's a common symptom of AI Psychosis according to the true believers in that assertion.

All my issues with its leader aside, great work OpenAI!

Re: GPT-6 Astra

#943

"""OpenAI called the model a "generational leap" for areas such as cybersecurity, professional work, software engineering, and science, with the company's president Greg Brockman saying that it could eventually be seen as the arrival of artificial general intelligence""" https://en.wikipedia.org/wiki/GPT-6_Astra according to Sam Altman from last year, an LLM should 'solve quantum gravity', if it is to count as AGI. D…

Bills are coming due, so marketing will change to fit the needs of the company's pocketbook. Please do not expect objectivity from Sam Altman or others with vested interests here. Its an impressive model, but AGI has always been a ridiculously nebulous concept used to inspire investors into giving away money.

Re: GPT-6 Astra

#944
post #679
post #640

- OpenAI claims Astra beats all benchmarks (compared to Fable and Opus, except "Humanity's Last Exam (w/ tools)"): https://openai.com/index/gpt-6-astra/ - Artificial Analysis scores Astra (max effort) as 61 points on intelligence, behind Opus 5. https://artificialanalysis.ai/models/gpt-6-astra Who is wrong here? Some benchmark results in Astra page for Fable and Opus are blank (-). What is Artificial Analysis intelli…

I really, really don't find the Artificial Analysis Intelligence Index credible anymore. It's some weighted score of benchmarks, and benchmarks increasingly don't reflect how good a model is. That should be obvious if you compare Gemini 3.8 Flash (which is an _excellent_ model especially for its price and TPS!! but 10min of prompting in any harness) will tell you it's nowhere near close to Sol/Astra. But AA scores Ge…

> but 10min of prompting in any harness) will tell you it's nowhere near close to Sol/Astra.

I code in both every day a lot and it is not obvious to me 3.8 is far behind

Re: GPT-6 Astra

#945
I'm sure this will be a great model.

Personally, I'm far away from screaming 'AGI is here!' from the rooftops, until jaggedness and silly mistakes disappear at the very least . (what is going on with that Mario Kart game...) So many benchmarks are 'best of x tries' or using very specific harnesses. AGI would not need a babysitter.

Honestly, even being able to do simple tasks like summarization or basic knowledge work without the constant paranoia of unforseen failure would be remarkable and useful.

Re: GPT-6 Astra

#946

I'm sure this will be a great model. Personally, I'm far away from screaming 'AGI is here!' from the rooftops, until jaggedness and silly mistakes disappear at the very least . (what is going on with that Mario Kart game...) So many benchmarks are 'best of x tries' or using very specific harnesses. AGI would not need a babysitter. Honestly, even being able to do simple tasks like summarization or basic knowledge work…

I'm still excited about this, Fable 5.1 and all the great open source models, but some very religious undertones have been entering the AI sphere

Re: GPT-6 Astra

#948
post #719

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

If you define AGI as "can do the work of a human sitting at a computer, end to end", then I'd say comparing yourself to it on a specific skill is the wrong test. Can you hand it a role and walk away for a day/week/month? I can’t yet. I think that I'd want at least two things it doesn't have: the ability to retain what it learned yesterday (without me carrying it in the context window and thus micromanaging it), and t…

I think fundamentally it is that. The ability to retain information.

Like given a specific task it can do a thing amazingly well, but can it recall a thing. Its memory seems like a giant filing cabinet and it has to go scan like 20 million tokens worth of memory to recover things previously talked about.

Human memory is more graph like, we don’t recall things exactly, but one thing links to another, we create a pattern of a thing, we mark what is important, and overtime what was important degrades or becomes less so.

I feel like what makes it lack intelligence is it never seems to learn. Like it kind of does, but then doesn’t persist once too many other things are learned.

I’m sure they’re probably working on this, but I feel like that is what I want far more than even better models, is a better memory system to recall and forget things that the models work on.

Re: GPT-6 Astra

#949

That hero video is interesting. A projector and speech. Maybe I'm in the minority here, but I find speech to text / text to speech (but not live audio mode) is quite comfortable and effective for coding now. The speech to text part can be frustrating if your local tts model does not have word match context for coding. Codex desktop does this remotely well but is slow. I've been experimenting with local software for m…

>this could bring us closer to the dream of more natural, social computing What I saw was multiple people living alone in a small box in a warehouse (probably filled with other boxes) with all of their natural, social interactions directed at a wall. I wonder if this is foreshadowing for the future of work, at least it is what work will look like as envisioned by OpenAI.

(I also wonder if the 1970s bit is actual and based on some well-known, "mother of all demos"-type of video or early experiment ...)
Post reply on HN