Live data from Hacker News

GPT-6 Astra

openai.com

241–250 of 1001 posts

Re: GPT-6 Astra

#241
post #59

Just two days ago, a preprint by Julia Stadlmann went up on arXiv [0] improving the prime gap from 246 to 240. Now OpenAI announces Astra has shown a gap of 186 [1]. That must really blow. [0] https://arxiv.org/abs/2608.31126 [1] https://cdn.openai.com/pdf/51126fac-1b68-4128-9666-c908bcc16...

Such a result should be considered worthless: the proof is 10MB of Lean. ( https://github.com/openai/PrimeGaps186 ). I can't think of a single mathematical proof being anywhere close to ten million characters. For all you know, 90% of the proof could be useless, 8% would be writing out Shakespeare, and 1% abusing another bug in Lean. Humanity gets zero value from that, aside from "some bot seems to think it's 186". U…

Terence Tao says something surprisingly similar in a recent talk (https://news.ycombinator.com/item?id=49056620 ) Not that the proof is worthless but that the value comes after it's revised into a cleanly understandable form and then canonicalized so that other mathematicians can use it.

Re: GPT-6 Astra

#242

This is wild: OpenAI is basically declaring that AGI is here. https://www.theverge.com/ai-artificial-intelligence/989601/o... “If we fast-forward a couple of years, and we look back and say, ‘When was it, really, that AGI was created?’ I think it’s going to be about this time, and I think it might be about this model,” OpenAI president Greg Brockman said during a Thursday press briefing. Later in the call, he added,…

Remember when the term "AGI" meant something? Pepperidge farm remembers

No, I really don't

Re: GPT-6 Astra

#243

Earlier quoted context omitted.

Such a result should be considered worthless: the proof is 10MB of Lean. ( https://github.com/openai/PrimeGaps186 ). I can't think of a single mathematical proof being anywhere close to ten million characters. For all you know, 90% of the proof could be useless, 8% would be writing out Shakespeare, and 1% abusing another bug in Lean. Humanity gets zero value from that, aside from "some bot seems to think it's 186". U…

Iirc some mainstream physycists never acknowledged quantum theory because they couldn’t accept that universe was that unintuitive and hard to understand. Ditto ones that opposed Einstein’s general relativity.

[flagged]

Re: GPT-6 Astra

#245
post #180

Earlier quoted context omitted.

I really wish it was called chain of instruction. Because it's definitely not thought.

They’re intermediate tokens, so I wish we called it what it is… ITG. The anthropomorphizing is out of control.

[deleted]

Re: GPT-6 Astra

#246
post #180

Earlier quoted context omitted.

I really wish it was called chain of instruction. Because it's definitely not thought.

They’re intermediate tokens, so I wish we called it what it is… ITG. The anthropomorphizing is out of control.

[deleted]

Re: GPT-6 Astra

#248

> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certa…

I really wish it was called chain of instruction. Because it's definitely not thought.

Yeah, basically they are using more computation to explore the solution space before producing the final answer.

Re: GPT-6 Astra

#249
post #160

Earlier quoted context omitted.

Such a result should be considered worthless: the proof is 10MB of Lean. ( https://github.com/openai/PrimeGaps186 ). I can't think of a single mathematical proof being anywhere close to ten million characters. For all you know, 90% of the proof could be useless, 8% would be writing out Shakespeare, and 1% abusing another bug in Lean. Humanity gets zero value from that, aside from "some bot seems to think it's 186". U…

Yeah. Unless human can verify it, not sure if it is certain or useful.

Wasn’t the proof of Fermatt’s Last Theorem proof similar in complexity?

Re: GPT-6 Astra

#250
post #82

ARC AGI-3 saturated by Astra! https://arcprize.org/leaderboard

But scored less on V2 and V1 ... too much overfitting?

https://mvakde.github.io/blog/44-on-arc-1/ makes a good case that all the performance on the arc agi tests is overfitting, based on the fact that v1 performance did not translate directly to v2 performance
Post reply on HN