Live data from Hacker News

GPT-6 Astra

openai.com

241–250 of 1001 posts

Re: GPT-6 Astra

#241

> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certa…

"OpenAI is pleased to announce our new model scores 85% on CreateTormentNexusBench - a >60% lead over our leading competitors!" Did someone get their "AI safety no-no list" and "Frontier features bingo card" mixed up, or did they just stop being able to tell the difference?

You joke, but a bunch of people here actually want that.

Re: GPT-6 Astra

#242
I remember when GPT-4 came out and the perceived performance upgrade seemed underwhelming for a major release compared to 3.5, especially how there were graphics going around showing the parameter size dwarfing the last model before it came out. It looked like we were past the perceivable differences from release to release that were immediately identifiable. Now the jump between 5 to 5.5 and 5.6 alone has changed how a lot of people approach AI, including me. Interested to see where it goes with 6.

Re: GPT-6 Astra

#244
post #58

Just two days ago, a preprint by Julia Stadlmann went up on arXiv [0] improving the prime gap from 246 to 240. Now OpenAI announces Astra has shown a gap of 186 [1]. That must really blow. [0] https://arxiv.org/abs/2608.31126 [1] https://cdn.openai.com/pdf/51126fac-1b68-4128-9666-c908bcc16...

Happened to multiple people I know.

Re: GPT-6 Astra

#245

I'm going to call it. By 2030 all software is done and complete. But we are going to have more and new jobs.

'all' software? aircraft flight control systems? infant heart monitors? drug manufacturing dose calibration controllers?

Re: GPT-6 Astra

#246
post #140

GPT 6 Astra benchmarks https://cdn.thenewstack.io/media/2026/09/358eb84a-screenshot... Performance is significantly higher than Fable 5.1 Source: https://thenewstack.io/openai-gpt6-astra-benchmarks/

> Performance is significantly higher than Fable 5.1 That's not clear. Need to see independent benchmarks first.

We need them pelicans on bikes.

Re: GPT-6 Astra

#247

What does 'Astra' here mean? Surely they must be referring to the Latin word. Because in another dead language of antiquity, Sanskrit, it means "weapon". Which would be a bit too on-the-nose.

Quite clearly in the same vein as Sol, Terra, Luna.

Re: GPT-6 Astra

#248
Argh! I hit a wrong keyboard shortcut and moved the entire thread.

Please stand by... it will all come back shortly

Re: GPT-6 Astra

#249
I want to take a step back: So, this is GPT-6 -- the natural number version release comparable to GPT-4 and GPT-5 from the past few years. The ARC-AGI-3 score is obviously impressive at 99.9% (we'll need to wait for more details on how they used the response API harness on GPT-6 Astra, wrt reasoning retention and compaction), but every other benchmarks seems to be a relatively modest improvement, comparable with any of the 'point' updates from AI labs.

If this is truly AGI (subject to one's definition of AGI still), then this is a very boring release of an AGI model. No video announcement, no presser, just a blog post (with some Twitter promo vids)?

As others mentioned, I'm starting to think OpenAI was under immense pressure to deliver an 'AGI' model for certain contractual reasons, but I never expected GPT-6 release to be this mundane and banal.

Re: GPT-6 Astra

#250

I wonder if this is going to be one of those days where you'll be like: Oh yeah I remember where I was when the first version of AGI launched

Define AGI first. The singularly ain't going to happen with LLMs.
Post reply on HN