GPT-6 Astra
811–820 of 1001 posts
Re: GPT-6 Astra
#812It's fun, but every new model release makes me even less interested to create cool stuff. Like, what's the point, if the next AI can do it in 5 seconds?
Then learn how to be creative and use the ai better?
Imagine Bob Ross prompting a robot what to paint...
Re: GPT-6 Astra
#813The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…
What a sad thing to say. These models are not even better than me at _writing code_, which is as well-suited a task for LLM agents as can possibly be, what with the structured environment and the exabytes of free annotated training data.
Of course, they are also not better than humans at writing, let alone at talking to my daughter, running a pathfinder campaign, decorating a room, being a therapist, etc.
Re: GPT-6 Astra
#814It's fun, but every new model release makes me even less interested to create cool stuff. Like, what's the point, if the next AI can do it in 5 seconds?
So this what a deflation economy is like. People don't want to do anything because they feel like whatever they do will be worthless in the near future. I've retreated to doing stuff with my hands. Wokdwork, DIY, that kind of thing. At least for now and the foreseeable future that doesn't seem pointless. Only problem is it's hard work yet nobody would pay me for it.
I might go more into table-tennis coaching, probably people would still prefer to be coached by a real person and not a robot for a very long time.
Re: GPT-6 Astra
#815It's fun, but every new model release makes me even less interested to create cool stuff. Like, what's the point, if the next AI can do it in 5 seconds?
I find it regularly incapable of doing the most basic things - maybe the idea is to think in terms of more detail? It’s kind of like saying what’s the point of poetry when a dictionary exists
I used GPT 5.6 Sol Extra High (Fast) for the last month, around 4-5h per day, and it managed to acomplished most of the tasks it had. It didn't really impress and often the end result needed one or two tweaks/fixes, but it did work and could create e2e solutions.
Re: GPT-6 Astra
#816The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…
Take it from the mouth of the creator of ARC-AGI: When we released ARC 3, I got asked, "when do you think a frontier model will saturate it?", and I answered "in about a year, though it depends on how much it gets explicitly targeted" That was 6 months ago, so the progress that Astra represents happened about 2x faster than I anticipated. I think the speed of progress will surprise a lot of people, and what the new m…
Re: GPT-6 Astra
#817Re: GPT-6 Astra
#818“allowing non-technical people to create and play custom games that go beyond rudimentary elements” Proceeds to generate the most generic, rudimentary, and unoriginal clone of Mario Kart
I think this kind of solves that. Or at least it is the start of it.
Most such games are kind of trivial. If people can easily just get AI to generate such games on the fly, then that'll hopefully be the end of predator pay-to-play games built on dark patterns.
Re: GPT-6 Astra
#819I want to take a step back: So, this is GPT-6 -- the natural number version release comparable to GPT-4 and GPT-5 from the past few years. The ARC-AGI-3 score is obviously impressive at 99.9% (we'll need to wait for more details on how they used the response API harness on GPT-6 Astra, wrt reasoning retention and compaction), but every other benchmarks seems to be a relatively modest improvement, comparable with any…
Does it learn? Does it experience? Can it connect with other agents, understand them, come to empathise with them and find a way to work with them better? The answer is no to all of these, and there are other problems as well. Yes, this model is trained to use a domain specific language to reason and plan over puzzle problems, and so it's programmers have cracked arc-agi-3 and that's a great achievement, but there is…
Re: GPT-6 Astra
#820Finally, OpenAI has a Fable/Mythos class model. 5.6 Sol felt like 5.5 on steroids, probably just a different checkpoint with a lot more RL post training. I wouldn't be surprised if there are some conceptual similarities to the kind of latent reasoning Anthropic sees in claude's J-space, although those aren't the same thing. Recurrent/looped transformers themselves aren't a new concept, but it's interesting to finally…
Sol easily outperforms Fable on every task I've tried it on.