Live data from Hacker News

GPT-6 Astra

openai.com

161–170 of 1001 posts

Re: GPT-6 Astra

#161
post #56

Just two days ago, a preprint by Julia Stadlmann went up on arXiv [0] improving the prime gap from 246 to 240. Now OpenAI announces Astra has shown a gap of 186 [1]. That must really blow. [0] https://arxiv.org/abs/2608.31126 [1] https://cdn.openai.com/pdf/51126fac-1b68-4128-9666-c908bcc16...

I think this builds straight upon her method, which she said could be improved herself so...

It cites to her at: [19] J. Stadlmann, On primes in arithmetic progressions and bounded gaps between many primes, Adv. Math. 468 (2025), Art. 110190. Numbered references use arXiv:2309.00425v3.

Though that's not her latest paper.

Re: GPT-6 Astra

#163

Finally, OpenAI has a Fable/Mythos class model. 5.6 Sol felt like 5.5 on steroids, probably just a different checkpoint with a lot more RL post training. I wouldn't be surprised if there are some conceptual similarities to the kind of latent reasoning Anthropic sees in claude's J-space, although those aren't the same thing. Recurrent/looped transformers themselves aren't a new concept, but it's interesting to finally…

yeah i'm wondering the same way... especially in light of the 20x debacle (where we found that 20x of Max vs 5x only applies to the 5hr limit, not the weekly limit, whereas OpenAI's 20x actually is 20x overall).

Also Opus 5 has been really tough to work with. I can't understand half of what it says, it's just so damn obscure.

Re: GPT-6 Astra

#164

GPT 6 Astra benchmarks https://cdn.thenewstack.io/media/2026/09/358eb84a-screenshot... Performance is significantly higher than Fable 5.1 Source: https://thenewstack.io/openai-gpt6-astra-benchmarks/

any benchmark where opus 5 achieves higher scores than fable 5 in any way is not a benchmark worth trusting.

Why would Anthropic trust and use these tests in their official comparisons?

Re: GPT-6 Astra

#165

You should know: AA index is only 61. Pretty surprised it’s that low.

This is actually a really good thing imo. If they didn't care about benchmaxxing it means that they really know that what they have in hand is good.

Re: GPT-6 Astra

#168
post #127

GPT 6 Astra benchmarks https://cdn.thenewstack.io/media/2026/09/358eb84a-screenshot... Performance is significantly higher than Fable 5.1 Source: https://thenewstack.io/openai-gpt6-astra-benchmarks/

> Performance is significantly higher than Fable 5.1 That's not clear. Need to see independent benchmarks first.

Artificial Analysis just published their aggregate score (61).

Still below Fable 5, let alone Fable 5.1.

EDIT: This is suspiciously low. Calls the relevance of existing benchmarks into question.

Re: GPT-6 Astra

#170

This is wild: OpenAI is basically declaring that AGI is here. https://www.theverge.com/ai-artificial-intelligence/989601/o... “If we fast-forward a couple of years, and we look back and say, ‘When was it, really, that AGI was created?’ I think it’s going to be about this time, and I think it might be about this model,” OpenAI president Greg Brockman said during a Thursday press briefing. Later in the call, he added,…

I hate the term "AGI" but IMO Fable, 5.6 Sol, et al. were already AGI.
Post reply on HN