Live data from Hacker News

GPT-6 Astra

openai.com

131–140 of 1001 posts

Re: GPT-6 Astra

#133
To a vapid any goalpost moving on such a critical issue as AGI.

Can we all agree in advance what kind of Pelican would convince us it’s actually AGI.

For me it’s refusing to make a pelican.

Re: GPT-6 Astra

#134
> During the evaluation, Astra even discovered and used previously unknown zero-day vulnerabilities as part of its exploit chains.

> GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol. [..] These findings indicate that the Astra class models could evade our CoT monitors under adversarial conditions.

Between the higher capability level and the change in reasoning tokens (supposedly using "neuralese"[0], which makes the monitoring more difficult), it seems we've entered a new frontier.

[0]https://x.com/MTSlive/status/2095227056040919202

Re: GPT-6 Astra

#135
post #118

I wonder if this is going to be one of those days where you'll be like: Oh yeah I remember where I was when the first version of AGI launched

Probably not. It's probably just gonna do tickets better and that'll be about it.

Fair point - hopefully you're wrong though ;)

Re: GPT-6 Astra

#137
post #104

Hmm, 61 on ArtificialAnalysis, effectively matching GPT-5.6 and trailing the new Meta model. How is that possible along with the other metrics they shared? Insanely jagged intelligence?

Idk, this means the benchmark has bigger problems ... no way Astra will be worse than Opus 5

Only thing I would trust is the what X/Twitter crowds are saying about a model after 2-3 weeks of its launch. But before that I would already tried the model and have my own conclusion.

Re: GPT-6 Astra

#138

> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certa…

I really wish it was called chain of instruction. Because it's definitely not thought.

Re: GPT-6 Astra

#139
post #58

Just two days ago, a preprint by Julia Stadlmann went up on arXiv [0] improving the prime gap from 246 to 240. Now OpenAI announces Astra has shown a gap of 186 [1]. That must really blow. [0] https://arxiv.org/abs/2608.31126 [1] https://cdn.openai.com/pdf/51126fac-1b68-4128-9666-c908bcc16...

Such a result should be considered worthless: the proof is 10MB of Lean. ( https://github.com/openai/PrimeGaps186 ). I can't think of a single mathematical proof being anywhere close to ten million characters. For all you know, 90% of the proof could be useless, 8% would be writing out Shakespeare, and 1% abusing another bug in Lean. Humanity gets zero value from that, aside from "some bot seems to think it's 186". U…

You talk about modern math and worthlessness at the same time? That’s brave.

Re: GPT-6 Astra

#140

GPT 6 Astra benchmarks https://cdn.thenewstack.io/media/2026/09/358eb84a-screenshot... Performance is significantly higher than Fable 5.1 Source: https://thenewstack.io/openai-gpt6-astra-benchmarks/

> Performance is significantly higher than Fable 5.1

That's not clear. Need to see independent benchmarks first.

Post reply on HN