I have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously? Even if I did trust an AI to get everything right, it's not like the AI can read my mind. If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really…
GPT-6 Astra
801–810 of 1001 posts
Re: GPT-6 Astra
#802Re: GPT-6 Astra
#803What does 'Astra' here mean? Surely they must be referring to the Latin word. Because in another dead language of antiquity, Sanskrit, it means "weapon". Which would be a bit too on-the-nose.
Seriously though the linear planetary naming is pleasant. I do hope the double meaning of "weapon", as you point out, is untrue.
Interestingly, Bruce Schnier just gave a fascinating talk about the Weapon aspect of AI, at DEFCON: https://youtu.be/eEBv0STiYhI very much worth watching
Re: GPT-6 Astra
#804Just two days ago, a preprint by Julia Stadlmann went up on arXiv [0] improving the prime gap from 246 to 240. Now OpenAI announces Astra has shown a gap of 186 [1]. That must really blow. [0] https://arxiv.org/abs/2608.31126 [1] https://cdn.openai.com/pdf/51126fac-1b68-4128-9666-c908bcc16...
Such a result should be considered worthless: the proof is 10MB of Lean. ( https://github.com/openai/PrimeGaps186 ). I can't think of a single mathematical proof being anywhere close to ten million characters. For all you know, 90% of the proof could be useless, 8% would be writing out Shakespeare, and 1% abusing another bug in Lean. Humanity gets zero value from that, aside from "some bot seems to think it's 186". U…
Re: GPT-6 Astra
#805They're just announcing later availability. No launch.
This is from Tibo on X.
Re: GPT-6 Astra
#806Where is the cure for cancer?
Would cancerbench be unethical?
Re: GPT-6 Astra
#807So what’s the tldr? :)
Does some things better
Rhetoric of AGI being tossed around, redefining the term
Re: GPT-6 Astra
#808The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…
I've personally been facing this lately 5.6 at max effort and fable have done tasks for me that I previously that would be a nearly 6mo project and it took me a week. It also did it better than I would have. The task was to build a high performance classification model. It not only helped make an entire data capture pipeline but also made the sythetic data basline needed. Then it proceeded to build and test 100 diffe…
How can you know it will have failed? I don't think it's that hard, if you clearly define the goal well, and have a bit more compute available, and do some intermediary bookkeeping.