Live data from Hacker News

GPT-6 Astra

openai.com

691–700 of 1001 posts

Re: GPT-6 Astra

#692

I have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously? Even if I did trust an AI to get everything right, it's not like the AI can read my mind. If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really…

I often feel like the use cases, demos, etc. that these Silicon Valley employees put out are based around their needs and how they operate. "Oh hey! Here's a demo of an AI planning out a 1-week trip to Paris!" No one in Middle America would just hand their credit card to an AI and let it come up with such a trip! I wish SV companies took more of the middle-class (and lower-middle-class) into consideration when coming…

I mean, not auto-purchasing with the card, no. But my wife has definitely used chatGPT to plan activities on the trips we've already booked.

Re: GPT-6 Astra

#693
post #324
post #77

ARC AGI-3 saturated by Astra! https://arcprize.org/leaderboard

The no-reasoning version scores 35% while the low reasoning one scores 17%? What?

I suspect this is "no reasoning set" which might be "default: medium" or perhaps some smart routing. I don't think it's literally "no reasoning".

Re: GPT-6 Astra

#694

I think the thing I'm most excited about is the increase in _user prompting_. If I give a poorly constrained/ambiguous prompt, I don't want the model one-shotting assumptions left and right. The demos of Fable/GPT-6 are impressive, but "real AGI" should act more like a collaborator than either a peon or overachiever. It's a tough balance to get right, and although this has been possible to achieve with additional pro…

>>> The demos of Fable/GPT-6 are impressive, but "real AGI" should act more like a collaborator than either a peon or overachiever.

I don't really agree. The thing that makes Fable feel like an actual collaborator is its ability to sus out your real intent when you give ambiguous instructions. It's really good at it.

I watched some reviews today and came way with the impression that Astra is not better than Sol in this regard. You still have to be very specific with your instructions. For example, you can say "why is it not committed yet?" and it will give you an explanation and say it's actually ready to be committed. But it won't commit unless you explicitly say so.

That sounds like a very tedious way of working with AI agents, but I understand some people want a high level of control.

Re: GPT-6 Astra

#695

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

> I am reasonably confident that there's essentially nothing that I am better than Fable at

Most things in the real world probably. I’m not saying AI can’t do it, but currently it’s bad. Try send an image of the inside of a broken toaster and how to fix. It’s laughable. Again, not saying AI will never do it, but am saying there are definitely large holes in knowledge.

Re: GPT-6 Astra

#696

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

Dumb single sample example: I asked Fable 5.1 to change from hard to soft deletion in an office map backend, and it used soft deletion for data which is synced from another system, but left hard deletion on for the mapping data itself (who sits where). For me it’s a pretty severe lack of judgement (like a red flag if I asked this in an interview).

Re: GPT-6 Astra

#697
> With Sites (opens in a new window) in ChatGPT, Astra can create, host, and share websites, web apps, and games directly from a prompt.

Oops, shots fired. A direct attack on the vibe coded app market. Replit, Lovable, etc.

Re: GPT-6 Astra

#698
post #497

I am really confused on how it can saturate ARC-AGI but still perform poorly on aggregated benchmarks: https://artificialanalysis.ai/models Perhaps if it was allowed this custom harness for all benchmarks it would similarily saturate?

Idk, they’re trying to sell a $500/mo/seat service to tell you what model is best. I think it’s in their interest to keep it confusing and opaque. Not exactly independent.

Re: GPT-6 Astra

#699
Sol has been very effective at schematic design (using Skidl) and at reviewing PCB layouts. But layout was still done manually by me. I'm very impressed and surprised to see they exactly a demo of Astra doing PCB layout. This is could be a game changer for electrial engineering! It already is since the schematic (and library management) is where a lot of the design work goes.

Re: GPT-6 Astra

#700
The model is probably excellent. The problem here is AGI having various definitions and many of them getting narrowed down to whatever makes benchmark numbers look good.
Post reply on HN