Live data from Hacker News

GPT-6 Astra

openai.com

121–130 of 1001 posts

Re: GPT-6 Astra

#123
post #58

Just two days ago, a preprint by Julia Stadlmann went up on arXiv [0] improving the prime gap from 246 to 240. Now OpenAI announces Astra has shown a gap of 186 [1]. That must really blow. [0] https://arxiv.org/abs/2608.31126 [1] https://cdn.openai.com/pdf/51126fac-1b68-4128-9666-c908bcc16...

I think this builds straight upon her method, which she said could be improved herself so...

Re: GPT-6 Astra

#125
Does anyone feel like everyone chasing the release of Anthropics Fabel 5.1 in a Mad Rush(tm)? In this situation it feels like tuning to benchmarks and other marketing devices feels like trusting Meta in mental health protection of users…

Re: GPT-6 Astra

#126

> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certa…

"OpenAI is pleased to announce our new model scores 85% on CreateTormentNexusBench - a >60% lead over our leading competitors!"

Did someone get their "AI safety no-no list" and "Frontier features bingo card" mixed up, or did they just stop being able to tell the difference?

Re: GPT-6 Astra

#127
post #92

The ARC-AGI-3 score is ridiculously high. Is this benchmaxxing or something way different? It's really hard to discern how we're approaching breakthroughs...

They explain why here: https://openai.com/index/how-two-settings-tripled-our-arc-ag... TL;DR all the other models are being crippled by limitations of their harness. >First, we noticed that after each game action, all private reasoning was discarded. This meant that with each action, GPT‑5.6 Sol was asked to figure out the game anew, unable to remember its past thinking. The model could still see a record of past mov…

Ok so the correct comparison would be to fix the harness on the old model and re-compare. Now they are comparing a new model to an old crippled one.

Re: GPT-6 Astra

#128
post #39

They're just announcing later availability. No launch.

Every frontier release nowadays is "we've launched*" * for a special group of customers that you're not in. Keep waiting peasant.

I mean tell Nvida to 100x their hardware output and you'll get what you want.
Post reply on HN