Live data from Hacker News

GPT-6 Astra

openai.com

851–860 of 1001 posts

Re: GPT-6 Astra

#851

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

> I've been at the point personally where I am reasonably confident that there's essentially nothing that I am better than Fable

What a sad thing to say. These models are not even better than me at _writing code_, which is as well-suited a task for LLM agents as can possibly be, what with the structured environment and the exabytes of free annotated training data.

Of course, they are also not better than humans at writing, let alone at talking to my daughter, running a pathfinder campaign, decorating a room, being a therapist, etc.

Re: GPT-6 Astra

#852
post #335

It's fun, but every new model release makes me even less interested to create cool stuff. Like, what's the point, if the next AI can do it in 5 seconds?

So this what a deflation economy is like. People don't want to do anything because they feel like whatever they do will be worthless in the near future. I've retreated to doing stuff with my hands. Wokdwork, DIY, that kind of thing. At least for now and the foreseeable future that doesn't seem pointless. Only problem is it's hard work yet nobody would pay me for it.

That's a good idea.

I might go more into table-tennis coaching, probably people would still prefer to be coached by a real person and not a robot for a very long time.

Re: GPT-6 Astra

#853
post #335

It's fun, but every new model release makes me even less interested to create cool stuff. Like, what's the point, if the next AI can do it in 5 seconds?

I find it regularly incapable of doing the most basic things - maybe the idea is to think in terms of more detail? It’s kind of like saying what’s the point of poetry when a dictionary exists

What model are you using?

I used GPT 5.6 Sol Extra High (Fast) for the last month, around 4-5h per day, and it managed to acomplished most of the tasks it had. It didn't really impress and often the end result needed one or two tweaks/fixes, but it did work and could create e2e solutions.

Re: GPT-6 Astra

#854
post #263

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

Take it from the mouth of the creator of ARC-AGI: When we released ARC 3, I got asked, "when do you think a frontier model will saturate it?", and I answered "in about a year, though it depends on how much it gets explicitly targeted" That was 6 months ago, so the progress that Astra represents happened about 2x faster than I anticipated. I think the speed of progress will surprise a lot of people, and what the new m…

Absolutely. You should never underestimate the compounding effect such a release can have. Having the right tools to create new tools.

Re: GPT-6 Astra

#856

“allowing non-technical people to create and play custom games that go beyond rudimentary elements” Proceeds to generate the most generic, rudimentary, and unoriginal clone of Mario Kart

Have you seen ads for mobile games, where there are seemingly 100 different versions of the same type of game (like tower defense types)? And they're all obviously the worst type of pay-to-play traps?

I think this kind of solves that. Or at least it is the start of it.

Most such games are kind of trivial. If people can easily just get AI to generate such games on the fly, then that'll hopefully be the end of predator pay-to-play games built on dark patterns.

Re: GPT-6 Astra

#857
post #823
post #227

I want to take a step back: So, this is GPT-6 -- the natural number version release comparable to GPT-4 and GPT-5 from the past few years. The ARC-AGI-3 score is obviously impressive at 99.9% (we'll need to wait for more details on how they used the response API harness on GPT-6 Astra, wrt reasoning retention and compaction), but every other benchmarks seems to be a relatively modest improvement, comparable with any…

Does it learn? Does it experience? Can it connect with other agents, understand them, come to empathise with them and find a way to work with them better? The answer is no to all of these, and there are other problems as well. Yes, this model is trained to use a domain specific language to reason and plan over puzzle problems, and so it's programmers have cracked arc-agi-3 and that's a great achievement, but there is…

Why does it need to empathise with something that doesn't have feelings in the first place? It clearly can learn from context. And experience? Again I don't see why it needs to feel anything.

Re: GPT-6 Astra

#858

Finally, OpenAI has a Fable/Mythos class model. 5.6 Sol felt like 5.5 on steroids, probably just a different checkpoint with a lot more RL post training. I wouldn't be surprised if there are some conceptual similarities to the kind of latent reasoning Anthropic sees in claude's J-space, although those aren't the same thing. Recurrent/looped transformers themselves aren't a new concept, but it's interesting to finally…

Sol easily outperforms Fable on every task I've tried it on.

what i found to work well with me is Fable for design / ideas and Sol for implementation. Codex models just tend to be more attentive and follow through instructions. Whereby claude models are weaker on this area (they tend to cut corners).

Re: GPT-6 Astra

#859

I'm going to call it. By 2030 all software is done and complete. But we are going to have more and new jobs.

> But we are going to have more and new jobs.

Like "fifth-rank junior assistant spouse in a comfort harem of an ultra-rich person".

Post reply on HN