I can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547 Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage and performance, more domains absorbed into the training distribution, and increasingly strong performance within that surface area. It seems more about coverage-driven competence. So…
GPT-6 Astra
491–500 of 1001 posts
Re: GPT-6 Astra
#492Re: GPT-6 Astra
#493I'm sure it's going to do great on all sorts of benchmarks, but the video--the actual marketing video that if anything is incentivised to overstate things--is full of careful cuts just before it would do anything that still wouldn't actually be that impressive. It's AGI, and it's going to upload photos, or change a background slide colour. Even the people hyping it up, who believe that it's really artificial intellig…
The games on mobile safari were broken. Buttons all misaligned in the kart racer one, the spaceship thing froze for a while, then kind of loaded but maybe not? Wasn't super compelling. I'm not trying to be too negative on it, it could be the best model right now, but it clearly isn't some agi god because things like that should have been caught (also should have been caught by human reviewers).
Re: GPT-6 Astra
#494The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…
Re: GPT-6 Astra
#495Re: GPT-6 Astra
#496Artificial analysis blog https://artificialanalysis.ai/articles/benchmarking-gpt-6-as...
Re: GPT-6 Astra
#497It's fun, but every new model release makes me even less interested to create cool stuff. Like, what's the point, if the next AI can do it in 5 seconds?
It's less lack of interest in creating that bothers me, it's my lack of interest in learning – it would surely be crazy for a SWE to care about how some new framework works anymore? Even if someone could reasonably argue that it might be slightly useful today there's almost zero chance it will be useful in 6-12 months times. But it's not just tech – my lack of interest in learning and creating is starting to generali…
The brain loves these kinds of shortcuts.
I don't need to think about the fine motor skills of hitting a baseball, it's just a motion now, and the game is still fun.
Re: GPT-6 Astra
#498The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…
Re: GPT-6 Astra
#499I have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously? Even if I did trust an AI to get everything right, it's not like the AI can read my mind. If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really…
What I desperately want is for 1password or stripe or even Google who already has much of my data, to o come up with a secure solution for online purchases with agentic credit cards where I can effectively get a phone prompt to authorize a purchase while the agent can fully own the checkout flow.
I have seen various things coming on the market for this, but none of them appear aimed at a consumer audience. And I am a firm believer at this point in keeping my payment authorization and history and credentials harness agnostic.
Re: GPT-6 Astra
#500The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…
Then realize LLMs have zero of what anyone would consider intelligence.