I think the thing I'm most excited about is the increase in _user prompting_. If I give a poorly constrained/ambiguous prompt, I don't want the model one-shotting assumptions left and right. The demos of Fable/GPT-6 are impressive, but "real AGI" should act more like a collaborator than either a peon or overachiever. It's a tough balance to get right, and although this has been possible to achieve with additional pro…
>>> The demos of Fable/GPT-6 are impressive, but "real AGI" should act more like a collaborator than either a peon or overachiever. I don't really agree. The thing that makes Fable feel like an actual collaborator is its ability to sus out your real intent when you give ambiguous instructions. It's really good at it. I watched some reviews today and came way with the impression that Astra is not better than Sol in th…
GPT-6 Astra
991–1000 of 1001 posts
Re: GPT-6 Astra
#992Finally, OpenAI has a Fable/Mythos class model. 5.6 Sol felt like 5.5 on steroids, probably just a different checkpoint with a lot more RL post training. I wouldn't be surprised if there are some conceptual similarities to the kind of latent reasoning Anthropic sees in claude's J-space, although those aren't the same thing. Recurrent/looped transformers themselves aren't a new concept, but it's interesting to finally…
> Canceling my Anthropic Max sub when this ships. At this point, it reads like people are cancelling old ones and getting new subscriptions every two to three days, whenever a new ,model drops, and quite possibly by the end of the week they are back to the old provider while still having active subscriptions with at least two to three others. Interesting times.
Re: GPT-6 Astra
#993I have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously? Even if I did trust an AI to get everything right, it's not like the AI can read my mind. If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really…
Most people go about their day absolutely minimizing the amount of mental energy they have to spend. They have priorities like kids, work, family, groceries, etc. If anything here can be automated its a fat win in their life. I have friends that are loving the features where meals are planned for them, food is home delivered, Uber is auto ordered, trip plans are made, etc. They don't mind paying more just to reduce mental load. Its a huhe market.
Re: GPT-6 Astra
#994I can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547 Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage and performance, more domains absorbed into the training distribution, and increasingly strong performance within that surface area. It seems more about coverage-driven competence. So…
You are conflating multiple things. 1) First, you are talking about positive forward transfer in continual learning. I've been giving talks for the past 6-7 years about how that community (I was one of the founders) went astray and wasn't focusing enough on that topic, but continual learning of the kind you are thinking isn't in any of these systems right now. I think some people left the Grok team to make a start-up…
Further, if there is a correlation, I'd bet it's not so much an intrinsic "creativity" trait, but more effectively higher creativity because more trials. That is, along the lines of Chollet's paper, a measure of creativity should be based on a fixed budget with fixed knowledge.
Among many other possibilities I haven't considered, perhaps another mechanism could be that because ADHD people spend more time thinking in less goal-oriented ways and mixing thoughts on accident, perhaps we do in fact gain some learned creativity via experience with vagueness[1]? But that might also imply that part of creativity is actually being able to diffuse more freely through thought space and lowering the barrier to attempted connections between ideas. That lower barrier leads to less likelihood of any "collision" being meaningful but maybe it's overcome by higher collision rates? Or maybe effectively higher order (not just pairwise) collisions?
Disclaimer in case it's not obvious: I don't know any of the literature on what creativity even means or how it's quantified.
[1] Which is me injecting an assumption that creativity ~= connecting things with no obvious or well-troden reasoning path between them.
edit -- oops just looked at your profile after seeing someone elses comment. I assume you are stating a fact then, leaving original anyway
Re: GPT-6 Astra
#995Finally, OpenAI has a Fable/Mythos class model. 5.6 Sol felt like 5.5 on steroids, probably just a different checkpoint with a lot more RL post training. I wouldn't be surprised if there are some conceptual similarities to the kind of latent reasoning Anthropic sees in claude's J-space, although those aren't the same thing. Recurrent/looped transformers themselves aren't a new concept, but it's interesting to finally…
> Canceling my Anthropic Max sub when this ships. At this point, it reads like people are cancelling old ones and getting new subscriptions every two to three days, whenever a new ,model drops, and quite possibly by the end of the week they are back to the old provider while still having active subscriptions with at least two to three others. Interesting times.
Re: GPT-6 Astra
#996It's fun, but every new model release makes me even less interested to create cool stuff. Like, what's the point, if the next AI can do it in 5 seconds?
Re: GPT-6 Astra
#997Re: GPT-6 Astra
#998I can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547 Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage and performance, more domains absorbed into the training distribution, and increasingly strong performance within that surface area. It seems more about coverage-driven competence. So…
Re: GPT-6 Astra
#999First impression: the model seems kind. Always intriguing