Live data from Hacker News

GPT-6 Astra

openai.com

791–800 of 1001 posts

Re: GPT-6 Astra

#791

I have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously? Even if I did trust an AI to get everything right, it's not like the AI can read my mind. If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really…

Despite access to """"""AGI""""""" all the marketing teams at these companies can only dream up 2 things, buying plane tickets and online shopping autonomously. Sometimes they're feeling extra spicy and throw in sorting emails or something along those lines. I suspect it's because it's tailored towards VCs and other similar rich ghouls as a replacement for their overworked and underpaid secretaries

They’re ridiculously facile use cases for what agentic AI can already do.

I don’t trust an agent with full executive power - yet - but I am essentially using GPT as my PA. Right now I’ve got it managing a construction project with recalcitrant contractors, an international move and visa tied to a property purchase, a short term holiday rental business, and basically just popping up in my life going “the situation is this, you need to do X/I suggest you send Y to Z, the email is prepped in your drafts”. I am of course also using it for software development, and have had it resolve every digital chore in my home life.

I guess it doesn’t make for a quick elevator pitch, but I’m finding it has reduced my cognitive overhead on a whole raft of fuckery that would otherwise have me shouting at inanimate objects.

Anyway, today I have to take the cat to the vet, email a bank, go and discuss ceiling systems, and review a proposed shopping list for a maintenance visit to a holiday rental. Can’t wait for the bleeding robots to get here.

Re: GPT-6 Astra

#792

I think the thing I'm most excited about is the increase in _user prompting_. If I give a poorly constrained/ambiguous prompt, I don't want the model one-shotting assumptions left and right. The demos of Fable/GPT-6 are impressive, but "real AGI" should act more like a collaborator than either a peon or overachiever. It's a tough balance to get right, and although this has been possible to achieve with additional pro…

>>> The demos of Fable/GPT-6 are impressive, but "real AGI" should act more like a collaborator than either a peon or overachiever. I don't really agree. The thing that makes Fable feel like an actual collaborator is its ability to sus out your real intent when you give ambiguous instructions. It's really good at it. I watched some reviews today and came way with the impression that Astra is not better than Sol in th…

> For example, you can say "why is it not committed yet?" and it will give you an explanation and say it's actually ready to be committed.

That's exactly what i want to happen. I hate when it assumes my direct question was an indirect instruction

Re: GPT-6 Astra

#793

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

In my view, intelligence includes an ability to learn and adapt to never-before-seen situations. And then general intelligence is an ability to apply that across a wide variety of domains. Machines can certainly recognize patterns and achieve goals through brute force trial and error. They can also use the results of previous iterations to change their behavior in future iterations, which we could call learning. I wo…

While you are mostly accurate in your definition, I’d argue we have discovered that intelligence is emergent/empirical not analytical. There is not a substrate we have yet to discern. Intelligence does not have to approach humanity to be AGI.

Re: GPT-6 Astra

#794

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

> I'd be curious to hear takes on what would make you think Astra is yet to be AGI

Give someone 10 remote employees for a few months, 5 of them human, 5 of them AI. After a few months, check to see if the humans (manager, other coworkers) can figure out who is AI and who isn't.

Would that be sufficient? I'd have to think about it. But AGI is supposed have human level capabilities, so this would be a necessary prerequisite.

None of the models are anywhere close to this.

Re: GPT-6 Astra

#795
post #227

I want to take a step back: So, this is GPT-6 -- the natural number version release comparable to GPT-4 and GPT-5 from the past few years. The ARC-AGI-3 score is obviously impressive at 99.9% (we'll need to wait for more details on how they used the response API harness on GPT-6 Astra, wrt reasoning retention and compaction), but every other benchmarks seems to be a relatively modest improvement, comparable with any…

Does it learn?

Does it experience?

Can it connect with other agents, understand them, come to empathise with them and find a way to work with them better?

The answer is no to all of these, and there are other problems as well. Yes, this model is trained to use a domain specific language to reason and plan over puzzle problems, and so it's programmers have cracked arc-agi-3 and that's a great achievement, but there is an asymmetry here. The arc team are well funded but are charged with providing a target for the vast ocean of funding, compute and talent everywhere else.

Most importantly, arc-agi-3 and the other benchmarks are all verifiable. The model can check if it's succeeded or not. They are not A* of course, but long horizon problems where you have to overcome minima to get the solution are not alien to AI either.

Re: GPT-6 Astra

#796

OpenAI is killing it now that they are more focused. Killing projects like Sora et al have seen it go from irrelevant to level footing with Anthropic. Sol is so much better than Fable 5. Then we get Astra (yet to use it) few days after Fable 5.1 (which is very impressive). Codex is slightly better than Claude Code. Good on Sam Altman getting back to basics and turning OpenAI around.

[dead]

Re: GPT-6 Astra

#797

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

> I'd be curious to hear takes on what would make you think Astra is yet to be AGI Give someone 10 remote employees for a few months, 5 of them human, 5 of them AI. After a few months, check to see if the humans (manager, other coworkers) can figure out who is AI and who isn't. Would that be sufficient? I'd have to think about it. But AGI is supposed have human level capabilities, so this would be a necessary prerequ…

I like it. Is the one who just absolutely ghosts everything on day 2 going to be human or AI?

Re: GPT-6 Astra

#799

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

They still seem pretty horrible at writing. Overly complicated prose, weird phrasing, poorly structured paragraphs. I don't know why they're so bad at communicating, but I feel very confident that humans are still much better at writing than any of these LLM models are, regardless of how advanced they are in other areas.

Some humans are much better at writing. Most humans are not. If you think they are, you are luckier than I am when it comes to the humans you need to communicate with.

As a non-native English speaker, I think the current LLMs write better English than me. I still write better than them in my native language (Norwegian), but the same cannot be said about most of my compatriots.

Re: GPT-6 Astra

#800

Lol their page finally loaded. They added an example scenario of "Filling in Form 1040" - which made me laugh out loud. That is indeed something most US citizens cannot accurately do even with expensive proprietary tax software services. Kind of a Hitchhiker's Guide to the Galaxy meme but where the tax code is so complicated we're implementing powerful AIs to be able to do it (hopefully) right.

You joke but...
Post reply on HN