Live data from Hacker News

GPT-6 Astra

openai.com

701–710 of 1001 posts

Re: GPT-6 Astra

#701

I have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously? Even if I did trust an AI to get everything right, it's not like the AI can read my mind. If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really…

I often feel like the use cases, demos, etc. that these Silicon Valley employees put out are based around their needs and how they operate.

"Oh hey! Here's a demo of an AI planning out a 1-week trip to Paris!" No one in Middle America would just hand their credit card to an AI and let it come up with such a trip!

I wish SV companies took more of the middle-class (and lower-middle-class) into consideration when coming up with such demos.

(Note: I live in SF)

Re: GPT-6 Astra

#704

I have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously? Even if I did trust an AI to get everything right, it's not like the AI can read my mind. If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really…

I often feel like the use cases, demos, etc. that these Silicon Valley employees put out are based around their needs and how they operate. "Oh hey! Here's a demo of an AI planning out a 1-week trip to Paris!" No one in Middle America would just hand their credit card to an AI and let it come up with such a trip! I wish SV companies took more of the middle-class (and lower-middle-class) into consideration when coming…

I mean, not auto-purchasing with the card, no. But my wife has definitely used chatGPT to plan activities on the trips we've already booked.

Re: GPT-6 Astra

#705
post #327
post #78

ARC AGI-3 saturated by Astra! https://arcprize.org/leaderboard

The no-reasoning version scores 35% while the low reasoning one scores 17%? What?

I suspect this is "no reasoning set" which might be "default: medium" or perhaps some smart routing. I don't think it's literally "no reasoning".

Re: GPT-6 Astra

#706

I think the thing I'm most excited about is the increase in _user prompting_. If I give a poorly constrained/ambiguous prompt, I don't want the model one-shotting assumptions left and right. The demos of Fable/GPT-6 are impressive, but "real AGI" should act more like a collaborator than either a peon or overachiever. It's a tough balance to get right, and although this has been possible to achieve with additional pro…

>>> The demos of Fable/GPT-6 are impressive, but "real AGI" should act more like a collaborator than either a peon or overachiever.

I don't really agree. The thing that makes Fable feel like an actual collaborator is its ability to sus out your real intent when you give ambiguous instructions. It's really good at it.

I watched some reviews today and came way with the impression that Astra is not better than Sol in this regard. You still have to be very specific with your instructions. For example, you can say "why is it not committed yet?" and it will give you an explanation and say it's actually ready to be committed. But it won't commit unless you explicitly say so.

That sounds like a very tedious way of working with AI agents, but I understand some people want a high level of control.

Re: GPT-6 Astra

#707

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

> I am reasonably confident that there's essentially nothing that I am better than Fable at

Most things in the real world probably. I’m not saying AI can’t do it, but currently it’s bad. Try send an image of the inside of a broken toaster and how to fix. It’s laughable. Again, not saying AI will never do it, but am saying there are definitely large holes in knowledge.

Re: GPT-6 Astra

#708

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

Dumb single sample example: I asked Fable 5.1 to change from hard to soft deletion in an office map backend, and it used soft deletion for data which is synced from another system, but left hard deletion on for the mapping data itself (who sits where). For me it’s a pretty severe lack of judgement (like a red flag if I asked this in an interview).

Re: GPT-6 Astra

#709
> With Sites (opens in a new window) in ChatGPT, Astra can create, host, and share websites, web apps, and games directly from a prompt.

Oops, shots fired. A direct attack on the vibe coded app market. Replit, Lovable, etc.

Re: GPT-6 Astra

#710
post #508

I am really confused on how it can saturate ARC-AGI but still perform poorly on aggregated benchmarks: https://artificialanalysis.ai/models Perhaps if it was allowed this custom harness for all benchmarks it would similarily saturate?

Idk, they’re trying to sell a $500/mo/seat service to tell you what model is best. I think it’s in their interest to keep it confusing and opaque. Not exactly independent.
Post reply on HN