Live data from Hacker News

GPT-6 Astra

openai.com

511–520 of 1001 posts

Re: GPT-6 Astra

#511
post #335

It's fun, but every new model release makes me even less interested to create cool stuff. Like, what's the point, if the next AI can do it in 5 seconds?

The point is to inject something into the process that these AIs can't do for you.

People SHOULD feel like making a useless Mario Kart clone isn't worth the effort anymore. They should, instead, be trying to figure out how to actually use these models to make something that doesn't feel like a useless Mario Kart clone.

Re: GPT-6 Astra

#512

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

Only someone who doesn't do any work of any meaningful difficulty could think these models have anything to do with AGI. Today I spent half a day trying to solve a moderately interesting software engineering problem. I was switching between GPT-5.6 Sol and Fable 5.1 to check each other's work in Cursor. And the result was gradually driving me insane. As the models struggled to find a solution that would actually work…

I still routinely have this experience too. But Sol and Fable feel closer and I have this experience less with them than with their predecessors.

Re: GPT-6 Astra

#513

OpenAI is killing it now that they are more focused. Killing projects like Sora et al have seen it go from irrelevant to level footing with Anthropic. Sol is so much better than Fable 5. Then we get Astra (yet to use it) few days after Fable 5.1 (which is very impressive). Codex is slightly better than Claude Code. Good on Sam Altman getting back to basics and turning OpenAI around.

Codex's lack of auto-mode is what prevents me from using it for serious work compared to Claude Code.

It has had automode for a bit now. I use it every day at work.

Re: GPT-6 Astra

#515
For people skeptical of AGI. Consider the following:

15 years ago if you were the sole proprietor of these models, would you be able to hold a dozen remote junior engineer jobs? Maybe even more? These models could certainly pass all interviews with flying colors and even survive independently in a company role.

I think sole ownership of AI 15 years ago could be worth north of $10 million per year. Just as rank-and-file employees.

Re: GPT-6 Astra

#517
I might be jaded, but these examples look silly, stereotyped, and absolutely how of touch with the nuances and the complexities of what real people would actually want/need to do in this specific situations.

Re: GPT-6 Astra

#518
post #335

It's fun, but every new model release makes me even less interested to create cool stuff. Like, what's the point, if the next AI can do it in 5 seconds?

Yeah. I’m almost glad I didn’t invest any time in any of my 100s ideas for a startup. Most of them would be destroyed by AI by now. But, you can create cool stuff just for yourself. That’s the upside. It’s just hard to make a living on cool stuff for yourself.

For some reason, I feel much less excited about creating things myself just knowing that ai can do it in 1/10th of the time. Even if I know it wouldn’t turn into a business or make me money. I don’t know why that is, but I was much more motivated to build anything (even things just for myself) before ai. Kinda depressing

Re: GPT-6 Astra

#519
post #360
post #335

It's fun, but every new model release makes me even less interested to create cool stuff. Like, what's the point, if the next AI can do it in 5 seconds?

Is there a point in playing Chess or Go when you know there's a computer out there that can beat you (and everyone else)?

You can play PvP in those games. Not really the same with developing software. In fact, not using ai would probably make you lose if there was some “software PvP” mode or development.

Re: GPT-6 Astra

#520
post #416

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

AGI to me is reached once the intelligence is self motivated, i.e. it doesn't rely on us prompting it into action. I don't see how LLMs will ever get to that stage.

Wouldn't agents that do inference in an infinite loop pass that bar?
Post reply on HN