I'm sure this will be a great model. Personally, I'm far away from screaming 'AGI is here!' from the rooftops, until jaggedness and silly mistakes disappear at the very least . (what is going on with that Mario Kart game...) So many benchmarks are 'best of x tries' or using very specific harnesses. AGI would not need a babysitter. Honestly, even being able to do simple tasks like summarization or basic knowledge work…
GPT-6 Astra
941–950 of 1001 posts
Re: GPT-6 Astra
#942Re: GPT-6 Astra
#943The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…
If you define AGI as "can do the work of a human sitting at a computer, end to end", then I'd say comparing yourself to it on a specific skill is the wrong test. Can you hand it a role and walk away for a day/week/month? I can’t yet. I think that I'd want at least two things it doesn't have: the ability to retain what it learned yesterday (without me carrying it in the context window and thus micromanaging it), and t…
Like given a specific task it can do a thing amazingly well, but can it recall a thing. Its memory seems like a giant filing cabinet and it has to go scan like 20 million tokens worth of memory to recover things previously talked about.
Human memory is more graph like, we don’t recall things exactly, but one thing links to another, we create a pattern of a thing, we mark what is important, and overtime what was important degrades or becomes less so.
I feel like what makes it lack intelligence is it never seems to learn. Like it kind of does, but then doesn’t persist once too many other things are learned.
I’m sure they’re probably working on this, but I feel like that is what I want far more than even better models, is a better memory system to recall and forget things that the models work on.
Re: GPT-6 Astra
#944That hero video is interesting. A projector and speech. Maybe I'm in the minority here, but I find speech to text / text to speech (but not live audio mode) is quite comfortable and effective for coding now. The speech to text part can be frustrating if your local tts model does not have word match context for coding. Codex desktop does this remotely well but is slow. I've been experimenting with local software for m…
>this could bring us closer to the dream of more natural, social computing What I saw was multiple people living alone in a small box in a warehouse (probably filled with other boxes) with all of their natural, social interactions directed at a wall. I wonder if this is foreshadowing for the future of work, at least it is what work will look like as envisioned by OpenAI.
Re: GPT-6 Astra
#945Re: GPT-6 Astra
#946I can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547 Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage and performance, more domains absorbed into the training distribution, and increasingly strong performance within that surface area. It seems more about coverage-driven competence. So…
You are conflating multiple things. 1) First, you are talking about positive forward transfer in continual learning. I've been giving talks for the past 6-7 years about how that community (I was one of the founders) went astray and wasn't focusing enough on that topic, but continual learning of the kind you are thinking isn't in any of these systems right now. I think some people left the Grok team to make a start-up…
Very strong reasoning here. Is there anything this ADHD condition cannot explain?
Re: GPT-6 Astra
#947Re: GPT-6 Astra
#948It's fun, but every new model release makes me even less interested to create cool stuff. Like, what's the point, if the next AI can do it in 5 seconds?
> Like, what's the point, if the next AI can do it in 5 seconds? I built a phone app recently, not released to the public, just an idea I had for ages but could never spend the time actually building. Its 100% vibe coded, and took me a few weekends to build... I'm talking a few hours in total. The point I'm making is that you now have the power to create stuff you would never have had the time to build. You can think…
Re: GPT-6 Astra
#949It's fun, but every new model release makes me even less interested to create cool stuff. Like, what's the point, if the next AI can do it in 5 seconds?
> Like, what's the point, if the next AI can do it in 5 seconds? Live a life doing whatever makes you happy. Post-work society is an inevitability if we don't destroy our planet.
we should be more concerned by post-wages society
Work will always exist regardless how useful it is
Re: GPT-6 Astra
#950I think the thing I'm most excited about is the increase in _user prompting_. If I give a poorly constrained/ambiguous prompt, I don't want the model one-shotting assumptions left and right. The demos of Fable/GPT-6 are impressive, but "real AGI" should act more like a collaborator than either a peon or overachiever. It's a tough balance to get right, and although this has been possible to achieve with additional pro…
That balance probably depends on the human, and the context. If you are a beginner in a field the model should not assume you know what you are doing. On the other hand, for an expert it should try to work out what you mean with your vaguely worded order. What I think should happen is that it should update its memory with notes on the proficiency level of the user, so it gets the balance right over time. This is a pr…
Which can also involve just asking for the users level of experience