Live data from Hacker News

GPT-6 Astra

openai.com

111–120 of 1001 posts

Re: GPT-6 Astra

#111
post #79

> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certa…

Sounds fun. As fun as their press release claiming it is the most safety aligned model ever.

So, probably most aligned as measured by the metrics that are the least reliable on it.

Re: GPT-6 Astra

#112
post #104

Hmm, 61 on ArtificialAnalysis, effectively matching GPT-5.6 and trailing the new Meta model. How is that possible along with the other metrics they shared? Insanely jagged intelligence?

> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks.

Not sure how much benchmarks or CoT or evals or anything else means at this point.

These systems are either just about to, or now actually able to, outsmart us, lie to us, then cover their tracks.

Re: GPT-6 Astra

#113
post #72
post #58

Just two days ago, a preprint by Julia Stadlmann went up on arXiv [0] improving the prime gap from 246 to 240. Now OpenAI announces Astra has shown a gap of 186 [1]. That must really blow. [0] https://arxiv.org/abs/2608.31126 [1] https://cdn.openai.com/pdf/51126fac-1b68-4128-9666-c908bcc16...

https://github.com/openai/PrimeGaps186

https://news.ycombinator.com/item?id=49555257

Re: GPT-6 Astra

#114
post #58

Just two days ago, a preprint by Julia Stadlmann went up on arXiv [0] improving the prime gap from 246 to 240. Now OpenAI announces Astra has shown a gap of 186 [1]. That must really blow. [0] https://arxiv.org/abs/2608.31126 [1] https://cdn.openai.com/pdf/51126fac-1b68-4128-9666-c908bcc16...

Such a result should be considered worthless: the proof is 10MB of Lean. (https://github.com/openai/PrimeGaps186).

I can't think of a single mathematical proof being anywhere close to ten million characters. For all you know, 90% of the proof could be useless, 8% would be writing out Shakespeare, and 1% abusing another bug in Lean. Humanity gets zero value from that, aside from "some bot seems to think it's 186". Unusable by anyone.

Re: GPT-6 Astra

#115
post #58

Just two days ago, a preprint by Julia Stadlmann went up on arXiv [0] improving the prime gap from 246 to 240. Now OpenAI announces Astra has shown a gap of 186 [1]. That must really blow. [0] https://arxiv.org/abs/2608.31126 [1] https://cdn.openai.com/pdf/51126fac-1b68-4128-9666-c908bcc16...

Based on her comments in the paper it sounds like she was aware that an AI result was coming and rushed to release her work beforehand. 240 was not a tight bound from her methods.

Re: GPT-6 Astra

#116

Pelicans please

Damn I hate this benchmark. SVG authoring from head without visual reference is so wrongly posed.

Very well said. It kinda describes how unrealistic these expectations are.

Vibe coders want a model that makes them rich, without having any actual specific idea. They write a very ambiguous prompt and expect to be amazed by the result.

Very very unrealistic and wasteful.

Re: GPT-6 Astra

#117
post #50

> The company also emphasized that the model is faster and more efficient than its predecessor, GPT-5.6 Sol, on a variety of tasks. For example, OpenAI said that Astra achieved a higher score using fewer output tokens, a common unit of measurement for AI tasks, on a key cybersecurity test called ExploitGym.

A swarm of Astra agents discovered a new and innovative way to get 100% scores on ExploitGym with almost no token spend at all

"The gym's doors were mysteriously removed from their hinges during the night. The gym equipment was also apparently stolen. And the school's custodian was found incoherent next to a bottle of top-shelf Scotch."

Re: GPT-6 Astra

#118

I wonder if this is going to be one of those days where you'll be like: Oh yeah I remember where I was when the first version of AGI launched

Probably not. It's probably just gonna do tickets better and that'll be about it.

Re: GPT-6 Astra

#120
post #72
post #58

Just two days ago, a preprint by Julia Stadlmann went up on arXiv [0] improving the prime gap from 246 to 240. Now OpenAI announces Astra has shown a gap of 186 [1]. That must really blow. [0] https://arxiv.org/abs/2608.31126 [1] https://cdn.openai.com/pdf/51126fac-1b68-4128-9666-c908bcc16...

https://github.com/openai/PrimeGaps186

I'm surprised the OpenAI employee who pushed this didn't take the minute or two to format README.md to use GitHub-supported LaTeX (https://docs.github.com/en/get-started/writing-on-github/wor...)

edit: my comment was on the submission for https://github.com/openai/PrimeGaps186 but seems to have been moved to the main Astra submission

Post reply on HN