Live data from Hacker News

GPT-6 Astra

openai.com

991–1000 of 1001 posts

Re: GPT-6 Astra

#992
post #967

I finally got access. Here are the pelicans! https://tools.simonwillison.net/markdown-svg-renderer?url=ht... The "max" one at the bottom took 4 minutes 2 seconds and cost 63.206 cents. For comparison, here those new Astra pelicans are in a grid with the GPT-5.6 pelicans: https://static.simonwillison.net/static/2026/gpt-6-and-5.6-p...

Love it! You should do a big table comparing pelicans for all the models you tested across all the big companies!

Re: GPT-6 Astra

#993
post #847

I feel that for some time now, the biggest constraint when working with models is not their intelligence, but their speed. It does not matter how smart the model is, it will make mistakes, because the instructions are ambiguous and new facts are found during implementation. The biggest problem I've had working with software developers has always been the lag between seeing the results and steering towards the right d…

I think speed is going to be important for a second reason - ever since I've started using LLM's a lot in my work I enjoy it much less. The main reason is that I ask the LLM something and have to wait because I know it's going to prompt me at random every few minutes. A lot of the day has become staring aimlessly at the screen. The prompts are too random and frequent for me to do something else in the interim. It's p…

This is why I switched to codex --yolo running in a container/vm. Now it does what it needs to do without bugging me and I can do other stuff. When it pings me I know there is something to review.

Yes I know they can escape containers, but that's not what I'm tasking them with.

Re: GPT-6 Astra

#994
post #847

I feel that for some time now, the biggest constraint when working with models is not their intelligence, but their speed. It does not matter how smart the model is, it will make mistakes, because the instructions are ambiguous and new facts are found during implementation. The biggest problem I've had working with software developers has always been the lag between seeing the results and steering towards the right d…

I'm sorry but no - output quality matters a lot more for me.

Just yesterday I tried to use Google antigravity to do a side project I've had on the back burner for 10 years now. Gemini flash is insanely fast - at first I was amazed at how quickly I was getting responses, and it seemed to hold it's own in technical discussion, although sycophancy is next level. But then when I actually let it do the coding part it was just drivel. I wouldn't even bother improving that code - like cleaning up after a lazy unskilled coworker - throw everything away and start over because the foundation is just leading in bad direction.

I spun up Astra on the same problem and although it was sluggish in comparison, and much more pedantic about irrelevant details - the feedback/pushback was actually meaningful. The implementation PoC also took tweaking but we got on the same page really fast.

Gemini Flash 3.8 was just producing garbage ultra fast, Astra could actually be steered into a direction I want and it provides valuable/insightful feedback.

I don't have infinite reading capacity/mental stamina - I would rather the model take it's time and let me see something high quality rather than get bombarded with garbage. If it can be faster that's great - but I'll always default to smarter model. The only exception is stupid trivial tasks like log analysis and similar.

Re: GPT-6 Astra

#995
post #967

I finally got access. Here are the pelicans! https://tools.simonwillison.net/markdown-svg-renderer?url=ht... The "max" one at the bottom took 4 minutes 2 seconds and cost 63.206 cents. For comparison, here those new Astra pelicans are in a grid with the GPT-5.6 pelicans: https://static.simonwillison.net/static/2026/gpt-6-and-5.6-p...

This is very sus, I got almost the same pelican with Fable 5.1.

Re: GPT-6 Astra

#996
post #333

It's fun, but every new model release makes me even less interested to create cool stuff. Like, what's the point, if the next AI can do it in 5 seconds?

> Like, what's the point, if the next AI can do it in 5 seconds? I built a phone app recently, not released to the public, just an idea I had for ages but could never spend the time actually building. Its 100% vibe coded, and took me a few weekends to build... I'm talking a few hours in total. The point I'm making is that you now have the power to create stuff you would never have had the time to build. You can think…

Yes! This is exactly how I feel.

I have fun ideas that are kind of complex and I can just experiment and see now with a couple spare hours on the weekend.

Trying to explain to my non software engineering friends that they could build any small app they have an idea for and they’re like I wouldn’t know how, and I’m like, you just ask and you can literally figure it out now so quick.

Maybe software engineers are the most excited and the most cynical.

Re: GPT-6 Astra

#997

I can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547 Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage and performance, more domains absorbed into the training distribution, and increasingly strong performance within that surface area. It seems more about coverage-driven competence. So…

You are conflating multiple things. 1) First, you are talking about positive forward transfer in continual learning. I've been giving talks for the past 6-7 years about how that community (I was one of the founders) went astray and wasn't focusing enough on that topic, but continual learning of the kind you are thinking isn't in any of these systems right now. I think some people left the Grok team to make a start-up…

> Given that this is HN, and a non-trivial number of us have ADHD, creativity is positively correlated with ADHD.

big pharma must LOVE people like you parroting the ADHD bullshit all the time

Re: GPT-6 Astra

#998
Seeing what others have already created using GPT-6, it seem to be a new stepping stone in capabilities and overall "intelligence". However, some other details I think is worth bringing up is that this model is 70% more token efficient than GPT-5.6 Sol and consuming 1/3 of the tokens compared to Sol (max) in the Codex [1].

I am just thinking loudly here but, it seems like even though Astra is pricier than Sol, you might actually get more usage out of it? I did some digging myself and looking at FrontierCode and DeepSWE, Astra (low) seem to perform better than Sol (medium) and on par with Luna (max) while being somewhat on the same price range to Sol? [2][3].

And now we have four models to chose from, each with their varied reasoning efforts: Astra, Sol, Terra and Luna. Personally, I feel like Terra have turned into this middle child in a weird spot that's neither the option as cheap model because Luna is, yet it is not an good option for complex tasks because Sol is already good at it.

For background, I use Luna (xhigh) daily, I think its a fantastic and underrated model. Especially Luna (max). It is way more capable than what it looks like, I think people underestimate it because OpenAI described it as "roughly corresponds to the nano model tier used in earlier GPT-5 families" [4]. I also like Luna because it barely consumes my weekly usage. Last week, it only ate ~15% of my weekly usage. So usage is not an issue anymore. I never have to worry. It may not be the fastest model because, well, it reasons as max effort, but it does the job way better than I expect. Also, considering how much one saves on the weekly usage, one can probably turn on "fast mode". Haven't done it myself though.

Also, another thing that caught my eyes is this:

> Historically, models have used compaction to summarize work during long sessions, such as when debugging complex issues or tackling large refactors. Each compaction can leave out details about why a fix failed or how a component behaves.

> In Codex, Astra can keep notes across context windows, preserving accumulated details without repeatedly compressing them into a single summary. Earlier context windows remain searchable, so Astra can find requirements or test results from previous messages and tool outputs—even if that information wasn’t captured in its notes. You can enable this experimental feature in your Codex config.toml [5].

I was curious about this, because I know Luna (max) spews out tokens which can trigger compaction quite often. If you go to the config reference [6] and search for "features.context_management.experimental_mode", you will find this:

> Enable experimental context management. Rather than repeatedly compressing context into a single summary, it uses notes and searchable history to preserve accumulated details.

This is a very interesting feature and perhaps very useful during long horizon work in a thread where the conversation context window grows and compacts often.

1. https://artificialanalysis.ai/articles/benchmarking-gpt-6-as...

2. https://deepswe.datacurve.ai, Astra (low) got 67% $2.19, Sol (medium) 61% $1.42 and Luna (max) 67% $0.61

3. https://cognition.com/frontiercode, Astra (low) 45.3% $1.60, Sol (medium) 39.9% $3.12, Luna (max) 39.8% $0.36

4. https://developers.openai.com/api/docs/models/gpt-5.6-luna

5. https://openai.com/index/gpt-6-astra

6. https://learn.chatgpt.com/docs/config-file/config-reference

Re: GPT-6 Astra

#999

its funny seeing HN commenters justifying how they are most certainly intelligent, but cant agree on what that is.

Its probably reached a certain level of intelligence but only operating in a very confined environment.

And I assume heavily language based.

Human brains use language for communication and other things but it isnt the only part of intelligence. There is also the intelligence of adapting and surviving in the world.

Until the AI can have a virtual environment or use the physical environment to interact with. I am still not convinced.

Post reply on HN