Live data from Hacker News

GPT-6 Astra

openai.com

981–990 of 1001 posts

Re: GPT-6 Astra

#981
post #975

I finally got access. Here are the pelicans! https://tools.simonwillison.net/markdown-svg-renderer?url=ht... The "max" one at the bottom took 4 minutes 2 seconds and cost 63.206 cents. For comparison, here those new Astra pelicans are in a grid with the GPT-5.6 pelicans: https://static.simonwillison.net/static/2026/gpt-6-and-5.6-p...

Technically on par with gemini flash 3.8, but I give Astra more points for style, and for breaking from the pack by not adding the headgear, and the fish in a basket.

Re: GPT-6 Astra

#982
post #975

I finally got access. Here are the pelicans! https://tools.simonwillison.net/markdown-svg-renderer?url=ht... The "max" one at the bottom took 4 minutes 2 seconds and cost 63.206 cents. For comparison, here those new Astra pelicans are in a grid with the GPT-5.6 pelicans: https://static.simonwillison.net/static/2026/gpt-6-and-5.6-p...

The "max" pelican looks very serious, almost as if it's determined to win the race!

Re: GPT-6 Astra

#983
post #975

I finally got access. Here are the pelicans! https://tools.simonwillison.net/markdown-svg-renderer?url=ht... The "max" one at the bottom took 4 minutes 2 seconds and cost 63.206 cents. For comparison, here those new Astra pelicans are in a grid with the GPT-5.6 pelicans: https://static.simonwillison.net/static/2026/gpt-6-and-5.6-p...

Technically on par with gemini flash 3.8, but I give Astra more points for style, and for breaking from the pack by not adding the headgear, and the fish in a basket.

It's a different prompt. It says his prompt was "Generate an SVG of a pelican riding a bicycle" but one I saw earlier this week specifically talked about the blue helmet and fish in a basket.

Re: GPT-6 Astra

#984
post #975

I finally got access. Here are the pelicans! https://tools.simonwillison.net/markdown-svg-renderer?url=ht... The "max" one at the bottom took 4 minutes 2 seconds and cost 63.206 cents. For comparison, here those new Astra pelicans are in a grid with the GPT-5.6 pelicans: https://static.simonwillison.net/static/2026/gpt-6-and-5.6-p...

The "max" pelican looks very serious, almost as if it's determined to win the race!

Yes! And do you get the impression, like I do, that model effort and pelican effort seem correlated?

Re: GPT-6 Astra

#985
post #975

I finally got access. Here are the pelicans! https://tools.simonwillison.net/markdown-svg-renderer?url=ht... The "max" one at the bottom took 4 minutes 2 seconds and cost 63.206 cents. For comparison, here those new Astra pelicans are in a grid with the GPT-5.6 pelicans: https://static.simonwillison.net/static/2026/gpt-6-and-5.6-p...

63 cents is quite cheap isn't it? If you compare vs Fable5.1 Max @ a whooping $3.30

Re: GPT-6 Astra

#986
post #975

I finally got access. Here are the pelicans! https://tools.simonwillison.net/markdown-svg-renderer?url=ht... The "max" one at the bottom took 4 minutes 2 seconds and cost 63.206 cents. For comparison, here those new Astra pelicans are in a grid with the GPT-5.6 pelicans: https://static.simonwillison.net/static/2026/gpt-6-and-5.6-p...

Pretty nice pelicans!

Re: GPT-6 Astra

#987
post #975

I finally got access. Here are the pelicans! https://tools.simonwillison.net/markdown-svg-renderer?url=ht... The "max" one at the bottom took 4 minutes 2 seconds and cost 63.206 cents. For comparison, here those new Astra pelicans are in a grid with the GPT-5.6 pelicans: https://static.simonwillison.net/static/2026/gpt-6-and-5.6-p...

The most interesting part to me is the bike. I don't think any of them would actually work, but the mistakes feel somewhat human.

Re: GPT-6 Astra

#988

I think the thing I'm most excited about is the increase in _user prompting_. If I give a poorly constrained/ambiguous prompt, I don't want the model one-shotting assumptions left and right. The demos of Fable/GPT-6 are impressive, but "real AGI" should act more like a collaborator than either a peon or overachiever. It's a tough balance to get right, and although this has been possible to achieve with additional pro…

>>> The demos of Fable/GPT-6 are impressive, but "real AGI" should act more like a collaborator than either a peon or overachiever. I don't really agree. The thing that makes Fable feel like an actual collaborator is its ability to sus out your real intent when you give ambiguous instructions. It's really good at it. I watched some reviews today and came way with the impression that Astra is not better than Sol in th…

> For example, you can say "why is it not committed yet?" and it will give you an explanation and say it's actually ready to be committed. But it won't commit unless you explicitly say so.

Well thank God, because that's the correct behavior. When you wanted to commit you can literally just say "commit" and nothing else.

Imagine if you asked it "you didn't delete my production database did you?"... And then it deletes your production database

Re: GPT-6 Astra

#989

I can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547 Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage and performance, more domains absorbed into the training distribution, and increasingly strong performance within that surface area. It seems more about coverage-driven competence. So…

I don't understand what you mean by novel intelligence or what people actually expect from these kinds of "Ai" but what novel intelligence can humans claim? Everything we know or learn is based on what someone else figured out. How are llms any different in that respect?

Re: GPT-6 Astra

#990

I can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547 Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage and performance, more domains absorbed into the training distribution, and increasingly strong performance within that surface area. It seems more about coverage-driven competence. So…

You are conflating multiple things. 1) First, you are talking about positive forward transfer in continual learning. I've been giving talks for the past 6-7 years about how that community (I was one of the founders) went astray and wasn't focusing enough on that topic, but continual learning of the kind you are thinking isn't in any of these systems right now. I think some people left the Grok team to make a start-up…

Appreciate the input! Responses below to your first two items:

1) I'm leaning more into a broader sense, which is that given the priors a system already possesses, how efficient can it acquire competence on a novel task? If I'm reading the point you make, you're focussing on continual learning right? If so, I'm not necessarily restricting my statement above to that.

Here's another reframing: How much of the benchmark improvements come from overwhelmingly large training distributions vs improving the models for adapting to things genuinely outside of it?

2) Excellent points about crystalized and fluid intelligence. Wouldn't the LLM scaling gains be a representation of crystallized capabilities? In regard to Gf, that is exactly what I am asking about. That is what seems to be lacking, Gf like adaptation under genuine novelty.

My concern is that it's increasingly difficult to tell of what looks like Gf like behavior is really coming from better adaption vs. having broad priors from the model's large learned distributions.

Post reply on HN