The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…
If you define AGI as "can do the work of a human sitting at a computer, end to end", then I'd say comparing yourself to it on a specific skill is the wrong test. Can you hand it a role and walk away for a day/week/month? I can’t yet. I think that I'd want at least two things it doesn't have: the ability to retain what it learned yesterday (without me carrying it in the context window and thus micromanaging it), and t…
GPT-6 Astra
781–790 of 1001 posts
Re: GPT-6 Astra
#782I can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547 Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage and performance, more domains absorbed into the training distribution, and increasingly strong performance within that surface area. It seems more about coverage-driven competence. So…
The more diverse stuff it knows, the easier it will be to learn something new.
Re: GPT-6 Astra
#783That hero video is interesting. A projector and speech. Maybe I'm in the minority here, but I find speech to text / text to speech (but not live audio mode) is quite comfortable and effective for coding now. The speech to text part can be frustrating if your local tts model does not have word match context for coding. Codex desktop does this remotely well but is slow. I've been experimenting with local software for m…
Re: GPT-6 Astra
#784Re: GPT-6 Astra
#785I'm sure it's going to do great on all sorts of benchmarks, but the video--the actual marketing video that if anything is incentivised to overstate things--is full of careful cuts just before it would do anything that still wouldn't actually be that impressive. It's AGI, and it's going to upload photos, or change a background slide colour. Even the people hyping it up, who believe that it's really artificial intellig…
Re: GPT-6 Astra
#786The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…
Machines can certainly recognize patterns and achieve goals through brute force trial and error. They can also use the results of previous iterations to change their behavior in future iterations, which we could call learning. I wouldn’t necessarily say they are good at brand new situations, but there has definitely been progress.
However, last I checked, a seemingly very intelligent LLM still struggles to play Chess at a basic level, let alone drive a robot or other non-language tasks. Its architecture and ability to learn seem a long way off from being general.
Vision models, being able to encompass language and much more, seem to me like a theoretically closer step to AGI. Yet, there is a lot more to the world than just what we can see.
On the other hand, in humans, vision certainly is not necessary for intelligence. So there is something more fundamental, neither vision nor language, that high levels of intelligence are based upon. Once we figure that out, I think we will be able to build AGI.
Re: GPT-6 Astra
#787I have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously? Even if I did trust an AI to get everything right, it's not like the AI can read my mind. If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really…
I often feel like the use cases, demos, etc. that these Silicon Valley employees put out are based around their needs and how they operate. "Oh hey! Here's a demo of an AI planning out a 1-week trip to Paris!" No one in Middle America would just hand their credit card to an AI and let it come up with such a trip! I wish SV companies took more of the middle-class (and lower-middle-class) into consideration when coming…
Re: GPT-6 Astra
#788The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…
In my view, intelligence includes an ability to learn and adapt to never-before-seen situations. And then general intelligence is an ability to apply that across a wide variety of domains. Machines can certainly recognize patterns and achieve goals through brute force trial and error. They can also use the results of previous iterations to change their behavior in future iterations, which we could call learning. I wo…
This is exactly what ARC AGI tests
> And then general intelligence is an ability to apply that across a wide variety of domains.
My experience with Fable is that it can certainly apply that in a wide variety of domains
> However, last I checked, a seemingly very intelligent LLM still struggles to play Chess at a basic level
People also struggle to play Chess at a basic level. They only succeed by studying the game for a long time. I will concede that humans can do this and LLMs generally cannot.
Re: GPT-6 Astra
#789Re: GPT-6 Astra
#790I have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously? Even if I did trust an AI to get everything right, it's not like the AI can read my mind. If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really…
Despite access to """"""AGI""""""" all the marketing teams at these companies can only dream up 2 things, buying plane tickets and online shopping autonomously. Sometimes they're feeling extra spicy and throw in sorting emails or something along those lines. I suspect it's because it's tailored towards VCs and other similar rich ghouls as a replacement for their overworked and underpaid secretaries
I don’t trust an agent with full executive power - yet - but I am essentially using GPT as my PA. Right now I’ve got it managing a construction project with recalcitrant contractors, an international move and visa tied to a property purchase, a short term holiday rental business, and basically just popping up in my life going “the situation is this, you need to do X/I suggest you send Y to Z, the email is prepped in your drafts”. I am of course also using it for software development, and have had it resolve every digital chore in my home life.
I guess it doesn’t make for a quick elevator pitch, but I’m finding it has reduced my cognitive overhead on a whole raft of fuckery that would otherwise have me shouting at inanimate objects.
Anyway, today I have to take the cat to the vet, email a bank, go and discuss ceiling systems, and review a proposed shopping list for a maintenance visit to a holiday rental. Can’t wait for the bleeding robots to get here.