Live data from Hacker News

GPT-6 Astra

openai.com

71–80 of 1001 posts

Re: GPT-6 Astra

#72
post #58

Just two days ago, a preprint by Julia Stadlmann went up on arXiv [0] improving the prime gap from 246 to 240. Now OpenAI announces Astra has shown a gap of 186 [1]. That must really blow. [0] https://arxiv.org/abs/2608.31126 [1] https://cdn.openai.com/pdf/51126fac-1b68-4128-9666-c908bcc16...

https://github.com/openai/PrimeGaps186

Re: GPT-6 Astra

#73
post #31

The ARCC-AGI-3 performance is absolutely incredible. The magnitude of change here is so high that I'm almost incredulous. Is this real? Did the benchmark get gamed?

They used a custom harness. It's not a one-to-one comparison.

Re: GPT-6 Astra

#75
post #45

I guess this "limited set of organizations" is just the standard now. It's just incredibly deflating to see my future as a second class citizen has already come

This has always been the case for people that have not had piles of money.

I mean do you get access to the best yachts?

To the top of the 5 star hotels?

To the best resorts?

To the best military equipment?

Hell, the best computer equipment has nearly always been out of reach of the average person.

Re: GPT-6 Astra

#76
Very nice to see that this is even more token efficient than Sol, when Fable 5.1 is less so than the already bloated token budget of Fable 5.

Re: GPT-6 Astra

#77
post #48

I was thinking about canceling my claude max sub after a few bad experiences. Kept hitting my usage limit, the quality of code seemed worse than Sol. This just made my decision. I'm moving to Codex Pro.

This is AGI now. Why are you spending any of your time looking at the "quality of code"?

Re: GPT-6 Astra

#78

> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certa…

Sounds fun. As fun as their press release claiming it is the most safety aligned model ever.

Re: GPT-6 Astra

#79
post #51

> GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol. We have performed significant investigations on the monitorability and controllability of GPT-6 Astra. We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the m…

I am also interesting knowing how they determined the model was sandbagging rather than just making a poor decision.

Also, this paragraph makes me wonder about all their stats on the exploitation and misalignment charts. If the model is that good at hiding "incriminating information" and sandbagging, are they sure its alignment is that?

Post reply on HN