I finally got access. Here are the pelicans! https://tools.simonwillison.net/markdown-svg-renderer?url=ht... The "max" one at the bottom took 4 minutes 2 seconds and cost 63.206 cents. For comparison, here those new Astra pelicans are in a grid with the GPT-5.6 pelicans: https://static.simonwillison.net/static/2026/gpt-6-and-5.6-p...
GPT-6 Astra
981–990 of 1001 posts
Re: GPT-6 Astra
#982I finally got access. Here are the pelicans! https://tools.simonwillison.net/markdown-svg-renderer?url=ht... The "max" one at the bottom took 4 minutes 2 seconds and cost 63.206 cents. For comparison, here those new Astra pelicans are in a grid with the GPT-5.6 pelicans: https://static.simonwillison.net/static/2026/gpt-6-and-5.6-p...
Technically on par with gemini flash 3.8, but I give Astra more points for style, and for breaking from the pack by not adding the headgear, and the fish in a basket.
Re: GPT-6 Astra
#983I finally got access. Here are the pelicans! https://tools.simonwillison.net/markdown-svg-renderer?url=ht... The "max" one at the bottom took 4 minutes 2 seconds and cost 63.206 cents. For comparison, here those new Astra pelicans are in a grid with the GPT-5.6 pelicans: https://static.simonwillison.net/static/2026/gpt-6-and-5.6-p...
The "max" pelican looks very serious, almost as if it's determined to win the race!
Re: GPT-6 Astra
#984I finally got access. Here are the pelicans! https://tools.simonwillison.net/markdown-svg-renderer?url=ht... The "max" one at the bottom took 4 minutes 2 seconds and cost 63.206 cents. For comparison, here those new Astra pelicans are in a grid with the GPT-5.6 pelicans: https://static.simonwillison.net/static/2026/gpt-6-and-5.6-p...
Re: GPT-6 Astra
#985I finally got access. Here are the pelicans! https://tools.simonwillison.net/markdown-svg-renderer?url=ht... The "max" one at the bottom took 4 minutes 2 seconds and cost 63.206 cents. For comparison, here those new Astra pelicans are in a grid with the GPT-5.6 pelicans: https://static.simonwillison.net/static/2026/gpt-6-and-5.6-p...
Re: GPT-6 Astra
#986I finally got access. Here are the pelicans! https://tools.simonwillison.net/markdown-svg-renderer?url=ht... The "max" one at the bottom took 4 minutes 2 seconds and cost 63.206 cents. For comparison, here those new Astra pelicans are in a grid with the GPT-5.6 pelicans: https://static.simonwillison.net/static/2026/gpt-6-and-5.6-p...
Re: GPT-6 Astra
#987I think the thing I'm most excited about is the increase in _user prompting_. If I give a poorly constrained/ambiguous prompt, I don't want the model one-shotting assumptions left and right. The demos of Fable/GPT-6 are impressive, but "real AGI" should act more like a collaborator than either a peon or overachiever. It's a tough balance to get right, and although this has been possible to achieve with additional pro…
>>> The demos of Fable/GPT-6 are impressive, but "real AGI" should act more like a collaborator than either a peon or overachiever. I don't really agree. The thing that makes Fable feel like an actual collaborator is its ability to sus out your real intent when you give ambiguous instructions. It's really good at it. I watched some reviews today and came way with the impression that Astra is not better than Sol in th…
Well thank God, because that's the correct behavior. When you wanted to commit you can literally just say "commit" and nothing else.
Imagine if you asked it "you didn't delete my production database did you?"... And then it deletes your production database
Re: GPT-6 Astra
#988I can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547 Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage and performance, more domains absorbed into the training distribution, and increasingly strong performance within that surface area. It seems more about coverage-driven competence. So…
Re: GPT-6 Astra
#989I can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547 Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage and performance, more domains absorbed into the training distribution, and increasingly strong performance within that surface area. It seems more about coverage-driven competence. So…
You are conflating multiple things. 1) First, you are talking about positive forward transfer in continual learning. I've been giving talks for the past 6-7 years about how that community (I was one of the founders) went astray and wasn't focusing enough on that topic, but continual learning of the kind you are thinking isn't in any of these systems right now. I think some people left the Grok team to make a start-up…
1) I'm leaning more into a broader sense, which is that given the priors a system already possesses, how efficient can it acquire competence on a novel task? If I'm reading the point you make, you're focussing on continual learning right? If so, I'm not necessarily restricting my statement above to that.
Here's another reframing: How much of the benchmark improvements come from overwhelmingly large training distributions vs improving the models for adapting to things genuinely outside of it?
2) Excellent points about crystalized and fluid intelligence. Wouldn't the LLM scaling gains be a representation of crystallized capabilities? In regard to Gf, that is exactly what I am asking about. That is what seems to be lacking, Gf like adaptation under genuine novelty.
My concern is that it's increasingly difficult to tell of what looks like Gf like behavior is really coming from better adaption vs. having broad priors from the model's large learned distributions.