Live data from Hacker News

GPT-6 Astra

openai.com

821–830 of 1001 posts

Re: GPT-6 Astra

#821
post #333

It's fun, but every new model release makes me even less interested to create cool stuff. Like, what's the point, if the next AI can do it in 5 seconds?

I find it regularly incapable of doing the most basic things - maybe the idea is to think in terms of more detail? It’s kind of like saying what’s the point of poetry when a dictionary exists

What model are you using?

I used GPT 5.6 Sol Extra High (Fast) for the last month, around 4-5h per day, and it managed to acomplished most of the tasks it had. It didn't really impress and often the end result needed one or two tweaks/fixes, but it did work and could create e2e solutions.

Re: GPT-6 Astra

#822
post #263

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

Take it from the mouth of the creator of ARC-AGI: When we released ARC 3, I got asked, "when do you think a frontier model will saturate it?", and I answered "in about a year, though it depends on how much it gets explicitly targeted" That was 6 months ago, so the progress that Astra represents happened about 2x faster than I anticipated. I think the speed of progress will surprise a lot of people, and what the new m…

Absolutely. You should never underestimate the compounding effect such a release can have. Having the right tools to create new tools.

Re: GPT-6 Astra

#824

“allowing non-technical people to create and play custom games that go beyond rudimentary elements” Proceeds to generate the most generic, rudimentary, and unoriginal clone of Mario Kart

Have you seen ads for mobile games, where there are seemingly 100 different versions of the same type of game (like tower defense types)? And they're all obviously the worst type of pay-to-play traps?

I think this kind of solves that. Or at least it is the start of it.

Most such games are kind of trivial. If people can easily just get AI to generate such games on the fly, then that'll hopefully be the end of predator pay-to-play games built on dark patterns.

Re: GPT-6 Astra

#825
post #791
post #227

I want to take a step back: So, this is GPT-6 -- the natural number version release comparable to GPT-4 and GPT-5 from the past few years. The ARC-AGI-3 score is obviously impressive at 99.9% (we'll need to wait for more details on how they used the response API harness on GPT-6 Astra, wrt reasoning retention and compaction), but every other benchmarks seems to be a relatively modest improvement, comparable with any…

Does it learn? Does it experience? Can it connect with other agents, understand them, come to empathise with them and find a way to work with them better? The answer is no to all of these, and there are other problems as well. Yes, this model is trained to use a domain specific language to reason and plan over puzzle problems, and so it's programmers have cracked arc-agi-3 and that's a great achievement, but there is…

Why does it need to empathise with something that doesn't have feelings in the first place? It clearly can learn from context. And experience? Again I don't see why it needs to feel anything.

Re: GPT-6 Astra

#826

Finally, OpenAI has a Fable/Mythos class model. 5.6 Sol felt like 5.5 on steroids, probably just a different checkpoint with a lot more RL post training. I wouldn't be surprised if there are some conceptual similarities to the kind of latent reasoning Anthropic sees in claude's J-space, although those aren't the same thing. Recurrent/looped transformers themselves aren't a new concept, but it's interesting to finally…

Sol easily outperforms Fable on every task I've tried it on.

what i found to work well with me is Fable for design / ideas and Sol for implementation. Codex models just tend to be more attentive and follow through instructions. Whereby claude models are weaker on this area (they tend to cut corners).

Re: GPT-6 Astra

#827

I'm going to call it. By 2030 all software is done and complete. But we are going to have more and new jobs.

> But we are going to have more and new jobs.

Like "fifth-rank junior assistant spouse in a comfort harem of an ultra-rich person".

Re: GPT-6 Astra

#830

I dropped my claude subscription a few months ago, though I kept some credits to do this and that with claude, thinking that claude might do better for some tasks. A few days ago they were all expired. It feels like it’s time to let claude go.

They're all complimentary. I've had double subs at the top tiers for a while now. Best of both worlds
Post reply on HN