Live data from Hacker News

GPT-6 Astra

openai.com

781–790 of 1001 posts

Re: GPT-6 Astra

#781

OpenAI is killing it now that they are more focused. Killing projects like Sora et al have seen it go from irrelevant to level footing with Anthropic. Sol is so much better than Fable 5. Then we get Astra (yet to use it) few days after Fable 5.1 (which is very impressive). Codex is slightly better than Claude Code. Good on Sam Altman getting back to basics and turning OpenAI around.

I think it mostly shows that there is no moat and the only advantage the U.S companies have over the Chinese is more compute. Qwen Max, Kimi K3, GLM 5.3 are really close to Opus/Sol/Fable/Astra and they are open weights.

From my experience with complex coding tasks (AI infra), I don't think these open weight models are close.

Re: GPT-6 Astra

#782

I have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously? Even if I did trust an AI to get everything right, it's not like the AI can read my mind. If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really…

Despite access to """"""AGI""""""" all the marketing teams at these companies can only dream up 2 things, buying plane tickets and online shopping autonomously. Sometimes they're feeling extra spicy and throw in sorting emails or something along those lines. I suspect it's because it's tailored towards VCs and other similar rich ghouls as a replacement for their overworked and underpaid secretaries

tarpit ideas

Re: GPT-6 Astra

#783

I have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously? Even if I did trust an AI to get everything right, it's not like the AI can read my mind. If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really…

I really wonder what the demo people are thinking as there are many great uses of LLMs in day-to-day lives but all we get is these 3 rehashed use cases. No wonder normal people think LLMs still can't do anything.

Re: GPT-6 Astra

#784

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

I’m so tired of every model release being touted as AGI or similar. Since GPT-2.

Re: GPT-6 Astra

#785
post #335

It's fun, but every new model release makes me even less interested to create cool stuff. Like, what's the point, if the next AI can do it in 5 seconds?

I find it regularly incapable of doing the most basic things - maybe the idea is to think in terms of more detail?

It’s kind of like saying what’s the point of poetry when a dictionary exists

Re: GPT-6 Astra

#787

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

Maybe I'm out of the loop, but wasn't AGI the full-on scifi version of AI, where the AI is a persistent, conscious entity? I don't see how task benchmark scores are relevant for that.

Re: GPT-6 Astra

#788
post #78

ARC AGI-3 saturated by Astra! https://arcprize.org/leaderboard

AGI means that the answers to these benchmarks are now accessible, and the defense capabilities of the testing organizations are now negligible.

Re: GPT-6 Astra

#789

I have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously? Even if I did trust an AI to get everything right, it's not like the AI can read my mind. If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really…

Despite access to """"""AGI""""""" all the marketing teams at these companies can only dream up 2 things, buying plane tickets and online shopping autonomously. Sometimes they're feeling extra spicy and throw in sorting emails or something along those lines. I suspect it's because it's tailored towards VCs and other similar rich ghouls as a replacement for their overworked and underpaid secretaries

All those stupid Bell Labs researchers not inventing Uber or Tinder. How come they didn't just build the obviously popular and profitable businesses that became possible once they invented the internet?

Re: GPT-6 Astra

#790

I have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously? Even if I did trust an AI to get everything right, it's not like the AI can read my mind. If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really…

Maybe we are not the real audience. I wonder if they are trying to convince the advertising industry/investors that in the future it won't be Google search that stands between the consumer and the product, but rather their LLM.
Post reply on HN