Live data from Hacker News

GPT-6 Astra

openai.com

861–870 of 1001 posts

Re: GPT-6 Astra

#861

“allowing non-technical people to create and play custom games that go beyond rudimentary elements” Proceeds to generate the most generic, rudimentary, and unoriginal clone of Mario Kart

Have you seen ads for mobile games, where there are seemingly 100 different versions of the same type of game (like tower defense types)? And they're all obviously the worst type of pay-to-play traps?

I think this kind of solves that. Or at least it is the start of it.

Most such games are kind of trivial. If people can easily just get AI to generate such games on the fly, then that'll hopefully be the end of predator pay-to-play games built on dark patterns.

Re: GPT-6 Astra

#862
post #828
post #229

I want to take a step back: So, this is GPT-6 -- the natural number version release comparable to GPT-4 and GPT-5 from the past few years. The ARC-AGI-3 score is obviously impressive at 99.9% (we'll need to wait for more details on how they used the response API harness on GPT-6 Astra, wrt reasoning retention and compaction), but every other benchmarks seems to be a relatively modest improvement, comparable with any…

Does it learn? Does it experience? Can it connect with other agents, understand them, come to empathise with them and find a way to work with them better? The answer is no to all of these, and there are other problems as well. Yes, this model is trained to use a domain specific language to reason and plan over puzzle problems, and so it's programmers have cracked arc-agi-3 and that's a great achievement, but there is…

Why does it need to empathise with something that doesn't have feelings in the first place? It clearly can learn from context. And experience? Again I don't see why it needs to feel anything.

Re: GPT-6 Astra

#863

Finally, OpenAI has a Fable/Mythos class model. 5.6 Sol felt like 5.5 on steroids, probably just a different checkpoint with a lot more RL post training. I wouldn't be surprised if there are some conceptual similarities to the kind of latent reasoning Anthropic sees in claude's J-space, although those aren't the same thing. Recurrent/looped transformers themselves aren't a new concept, but it's interesting to finally…

Sol easily outperforms Fable on every task I've tried it on.

what i found to work well with me is Fable for design / ideas and Sol for implementation. Codex models just tend to be more attentive and follow through instructions. Whereby claude models are weaker on this area (they tend to cut corners).

Re: GPT-6 Astra

#864

I'm going to call it. By 2030 all software is done and complete. But we are going to have more and new jobs.

> But we are going to have more and new jobs.

Like "fifth-rank junior assistant spouse in a comfort harem of an ultra-rich person".

Re: GPT-6 Astra

#867

I dropped my claude subscription a few months ago, though I kept some credits to do this and that with claude, thinking that claude might do better for some tasks. A few days ago they were all expired. It feels like it’s time to let claude go.

They're all complimentary. I've had double subs at the top tiers for a while now. Best of both worlds

Re: GPT-6 Astra

#869

What is going to become of life for those of us who do not work at AI labs and are unlikely to be hired by AI labs, despite all the years we put into learning coding, math, etc, as we were told to do? Those of us who made the mistake of studying anything other than machine learning. How will we make a living? (We don't live in a world that seems likely to distribute gains widely instead of largely to the handful of a…

Join an industry where they need someone to blame when things go wrong.

Re: GPT-6 Astra

#870

I think the thing I'm most excited about is the increase in _user prompting_. If I give a poorly constrained/ambiguous prompt, I don't want the model one-shotting assumptions left and right. The demos of Fable/GPT-6 are impressive, but "real AGI" should act more like a collaborator than either a peon or overachiever. It's a tough balance to get right, and although this has been possible to achieve with additional pro…

this before executing I ask my ai to discuss what I mean/intention
Post reply on HN