It's fun, but every new model release makes me even less interested to create cool stuff. Like, what's the point, if the next AI can do it in 5 seconds?
> Like, what's the point, if the next AI can do it in 5 seconds? Live a life doing whatever makes you happy. Post-work society is an inevitability if we don't destroy our planet.
GPT-6 Astra
741–750 of 1001 posts
Re: GPT-6 Astra
#742The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…
If you define AGI as "can do the work of a human sitting at a computer, end to end", then I'd say comparing yourself to it on a specific skill is the wrong test. Can you hand it a role and walk away for a day/week/month? I can’t yet. I think that I'd want at least two things it doesn't have: the ability to retain what it learned yesterday (without me carrying it in the context window and thus micromanaging it), and t…
So, for now, humans need to stay in the loop and do low skilled labor to keep the skilled work the models do from going off the rails.
Re: GPT-6 Astra
#743I have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously? Even if I did trust an AI to get everything right, it's not like the AI can read my mind. If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really…
That’s exactly the problem I have with all this agent ideas too. Imagine you had a human concierge that is just waiting for your instructions and is as smart or a bit smarter than you. Would you just tell them “plan this holiday for me” or “order this food”? I don’t even trust my friends to get this right, why would I give this to someone else?
When I tell a bot to find the best value per volume for a reasonable quantity of unscented Dawn dish soap [so I can buy that], then: It often makes a complete mess of this seemingly-simple operation.
(And yeah, that is an actual thing that I've tried to accomplish with voice commands while standing in my kitchen and doing some dishes. It seems very simple, and it did not go well.
Maybe when we get the basics figured out we can start worrying about how inept it is at doing vacation planning.
It seems that this kind of thing isn't sorted at all, and that this is a very real problem for those who are in the bot business: These missed opportunities leave money on the table.)
Re: GPT-6 Astra
#744I have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously? Even if I did trust an AI to get everything right, it's not like the AI can read my mind. If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really…
Despite access to """"""AGI""""""" all the marketing teams at these companies can only dream up 2 things, buying plane tickets and online shopping autonomously. Sometimes they're feeling extra spicy and throw in sorting emails or something along those lines. I suspect it's because it's tailored towards VCs and other similar rich ghouls as a replacement for their overworked and underpaid secretaries
As for public releases: I wonder if it's because these examples are easy to relate to. Many websites are just a long tail of industry or use-case specific stuff. What's valuable to me probably means nothing to you. This is unlikely to resonate with people-wit-large (and LLMs are marketed broadly) or requires the reader to think (and marketing that requires thinking is bad these days).
Second, it's arguably a good litmus test. If it still can't do the worn out examples of plane tickets and shopping, which would be a good assumption since we've been demo'd these use-cases for 2 years at this point, then ...
Re: GPT-6 Astra
#745Re: GPT-6 Astra
#746The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…
> I am reasonably confident that there's essentially nothing that I am better than Fable at despite generally being substantively above average on human benchmarks I genuinely don’t understand how an adult can say this with a straight face. I can take any single of my hobbies, start a mildly advanced conversation with Fable about the hobby and, within 5-10 turns, get it to contradict itself about something fundamenta…
Your observations expose the brittleness of the benchmarks being used for Fable, where the 'reasonable confident' claimant is working in the problem space of those two benchmarks, and you're demonstrating the failure at depth of the very capability facade for the model.
Sure, both the model and the confident human get the first layer right in some benchmark, but at least the model, and probably the human as well reach a collapsing probability of accuracy quite quickly. They may be completely and unconditionally right that Fable is more intelligent than they are, but that has little relevance at intelligence in depth as measured against your hobbies and the risk of unsafe or false responses.
Being smarter than a select group of people under test conditions is not conclusive regarding AGI. Moving the goalposts to declare AGI via a shallow benchmark, when a model could be proven wrong by almost anyone with basic competency given a few turns of iterative depth is where the 'grift' resides.
We're so far from anything approaching actual AGI, and it is highly speculative to infer that the current approach to Machine Learning as applied in LLM is even on a path that leads to AGI. But sales and promotions teams gotta pose, and we can all hate both the game and the players.
Re: GPT-6 Astra
#747It's fun, but every new model release makes me even less interested to create cool stuff. Like, what's the point, if the next AI can do it in 5 seconds?
Re: GPT-6 Astra
#748The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…
I think I have the following questions about what AGI would look like:
1. Do you expect an AGI to be able to competently do any knowledge work that an able human does today?
I think this is implied by the "General" component. I would assume that anything we call AGI would be able to do any of these tasks if given the time and reference material needed.
2. How would you expect AGI to handle edge cases? (Missing context, no known solution, under specified instructions, over specified instructions)
I would expect an agent to be able to look at the context that work exists in and correctly attenuate it's intentions for these goals. Simpler solutions, more thorough reporting, etc based on the need.
3. In my work I attend meetings, write reports, write code, research things, etc. Would AGI be able to reliably do that?
I would say that AGI would need to do this. I would classify this as the "Intelligence" component. Obtaining context, building a model of a problem, solving it, and convincing others.
4. Would it be able to inspire trust in itself? Trust can be established through verification of it's outputs, the construction of introspective tools, no hallucinations, etc.
I would say yes to this as well. It would be a component of the "Intelligence" to know that buy in is more important than the completion of a task.
To these points, will Astra be able to do these things? If not, I would hesitate to call it AGI.
Re: GPT-6 Astra
#749OpenAI is killing it now that they are more focused. Killing projects like Sora et al have seen it go from irrelevant to level footing with Anthropic. Sol is so much better than Fable 5. Then we get Astra (yet to use it) few days after Fable 5.1 (which is very impressive). Codex is slightly better than Claude Code. Good on Sam Altman getting back to basics and turning OpenAI around.
Killing Sora was one of the worst mistakes they ever made
That announcement is when I stopped paying attention to them.
Re: GPT-6 Astra
#750Pelicans please