Live data from Hacker News

GPT-6 Astra

openai.com

931–940 of 1001 posts

Re: GPT-6 Astra

#931

I think the thing I'm most excited about is the increase in _user prompting_. If I give a poorly constrained/ambiguous prompt, I don't want the model one-shotting assumptions left and right. The demos of Fable/GPT-6 are impressive, but "real AGI" should act more like a collaborator than either a peon or overachiever. It's a tough balance to get right, and although this has been possible to achieve with additional pro…

>>> The demos of Fable/GPT-6 are impressive, but "real AGI" should act more like a collaborator than either a peon or overachiever. I don't really agree. The thing that makes Fable feel like an actual collaborator is its ability to sus out your real intent when you give ambiguous instructions. It's really good at it. I watched some reviews today and came way with the impression that Astra is not better than Sol in th…

Why would I ask a LLM a rhetorical question

Re: GPT-6 Astra

#932

I can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547 Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage and performance, more domains absorbed into the training distribution, and increasingly strong performance within that surface area. It seems more about coverage-driven competence. So…

You are conflating multiple things. 1) First, you are talking about positive forward transfer in continual learning. I've been giving talks for the past 6-7 years about how that community (I was one of the founders) went astray and wasn't focusing enough on that topic, but continual learning of the kind you are thinking isn't in any of these systems right now. I think some people left the Grok team to make a start-up…

What have you defined as creativity and intelligence?

Re: GPT-6 Astra

#933
post #641

- OpenAI claims Astra beats all benchmarks (compared to Fable and Opus, except "Humanity's Last Exam (w/ tools)"): https://openai.com/index/gpt-6-astra/ - Artificial Analysis scores Astra (max effort) as 61 points on intelligence, behind Opus 5. https://artificialanalysis.ai/models/gpt-6-astra Who is wrong here? Some benchmark results in Astra page for Fable and Opus are blank (-). What is Artificial Analysis intelli…

> Humanity's Last Exam (w/ tools) This is one of the only benchmarks that actually matters for testing the frontier however. Other benchmarks can be gamed by simply being more persistent, but HLE is a diverse set of open-ended research-level questions. It tests domain knowledge and problem solving skills. Burning more reasoning tokens may help somewhat but not as much as e.g. coding benchmarks.

Do we even know what the true ceiling of hle is? I'm pretty sure some of their public sample questions are wrong (ambiguous/nonsensical with most logical interpretation trivial)

Re: GPT-6 Astra

#934

I firmly believed OpenAI would take back the lead. It's healthy to have competition with Anthropic and, hopefully, other LLM providers. I'm excited to test their model. Their attitude towards developers/builders has been nice and appreciated for the last couple of months.

I find that when someone comes from an attitude of abuse and hostility towards it users, then later without taking responsibility and owning its original actions says “we changed were being nice now” that it isn’t really a change but a farce to continue to trick you. I used to be more gullible and naive and fall for these tricks, but not anymore.

You change when you take ownership of your past mistakes and share with those you harmed how you have done that. With some reparation and restoration of past harms. I haven’t seen openAI do that.

Re: GPT-6 Astra

#935
Please, I have a question: in September 2026, what is the price differential between buying GPT inference tokens vs. a $20/month subscription?

I have moved to just using Google Gemini occasionally paying for tokens, no subscription, and it is OK, given my low level of use.

I would like to add gpt-6 astra to my toolbox, but I probably would use it infrequently.

Re: GPT-6 Astra

#936
post #862

I feel that for some time now, the biggest constraint when working with models is not their intelligence, but their speed. It does not matter how smart the model is, it will make mistakes, because the instructions are ambiguous and new facts are found during implementation. The biggest problem I've had working with software developers has always been the lag between seeing the results and steering towards the right d…

> It does not matter how smart the model is, it will make mistakes It does matter, otherwise why are we all using GPT 5.6 rather than GPT 3.5? Because it's way smarter, makes less mistakes and therefore finishes tasks faster. The smarter the model is, the faster it can complete what you actually wanted. > Regardless, working on the wrong things is time wasted Agreed. And "smart" for me, would mean understanding what…

[flagged]

Re: GPT-6 Astra

#937

Please, I have a question: in September 2026, what is the price differential between buying GPT inference tokens vs. a $20/month subscription? I have moved to just using Google Gemini occasionally paying for tokens, no subscription, and it is OK, given my low level of use. I would like to add gpt-6 astra to my toolbox, but I probably would use it infrequently.

$20 per month should get you far enough if you use luna, it's a decent model. If you plan to try astra you need to try make it use as little context as possible to protect your usage limit - or bank a reset mid task to allow it to keep going.

If you use a decent part of a subscription - of any tier - you could be saving a lot of money. According to YT that could even be ~80-90% less than paying for tokens but take that with a grain of salt.

Post reply on HN