> Starting today, GPT‑5.6 Luna, our fastest and most affordable model, will cost 80% less, I don't have the words. I genuinely thought we were in a stage where we were plateauing and going in for 5-10% improvements over months. Seeing spikes like this makes me question about where the floor really is.
Advancing the price-performance frontier with GPT‑5.6
171–180 of 424 posts
Re: Advancing the price-performance frontier with GPT‑5.6
#172Earlier quoted context omitted.
Experienced similar between 5.4-mini vs 5.6-luna in our own pipelines but after spending some time on prompt optimization and testing out various reasoning effort levels 5.6-luna was well worth it. Did you just replace model selection while keeping everything else in place or spend some time on evaling with newer prompts etc?
No we kept prompts as is, just swapped model. The prompt is already quite optimized for the task. How would updating it possibly make a more intelligent model spend less tokens than a less intelligent model? Care to elaborate?
Re: Advancing the price-performance frontier with GPT‑5.6
#173"Half the money I spend on advertising is wasted; the trouble is I don't know which half." -John Wanamaker This applies even more strongly to model choosing. I know for a fact that majority of my work doesn't require a very strong model, but separating the trivial and non-trivial tasks is a famously hard problem (if at all decidable).
Re: Advancing the price-performance frontier with GPT‑5.6
#174Earlier quoted context omitted.
Burning the weights into silicon would be many orders of magnitude increase, not just 10x. It's kind of crazy that this hockey stick the AI hype bros talk about seems more and more every day like it might be real
And don't underestimate how much money Google, Microsoft, Amazon and Meta still have to spend on this tech. Blocking Fable for sure made it very politicl a lot sooner than i expected it to happen. and because China already has massive problems of getting access, they are pushing it on hardware too like what Huawai did without EUV. It seems China is already able to do DUV a lot sooner than others expected.
That's the media and in particular US KOLs of all sorts driving the wrong impression of China and other places. China and many other places for example have fast public transport that the US doesn't and can't even imagine today. They're not behind.
China's DUV still isn't that production grade (mass produce-able) so don't get that hyped up the wrong way (in a different direction).
The whole China-is-behind with tech and in particular semi wasn't that they can't. The truth is they spent decades in internal politics and corruption. That all got solved with the bans, so thank the bans! Jensen even said the bans were bad.
Re: Advancing the price-performance frontier with GPT‑5.6
#175Edit: Yes, 80% minus is still milking. Because you empower these greedy mega-corporations. Just look at the RAM prices increase, then you see that the more money you give these hungry dragons, they more they will eat up. Don't get fooled by their "less cost now" advertisement.
Re: Advancing the price-performance frontier with GPT‑5.6
#176Re: Advancing the price-performance frontier with GPT‑5.6
#177Earlier quoted context omitted.
It's marketing hyperbole, but Luna is more intelligent per dollar than deepseek-v4 pro. Cost means nothing without the associated capability
is it? i don't know how we measure these things, but here's one measurement that says v4 pro is better than luna: https://artificialanalysis.ai/models/comparisons/gpt-5-6-lun... presumably it's a much bigger model
Re: Advancing the price-performance frontier with GPT‑5.6
#178> Starting today, GPT‑5.6 Luna, our fastest and most affordable model, will cost 80% less, I don't have the words. I genuinely thought we were in a stage where we were plateauing and going in for 5-10% improvements over months. Seeing spikes like this makes me question about where the floor really is.
Hard to believe numbers. I don't mean that as a critique, but literally I am so impressed. Even if the model is a few percent lower for performance but is 80+% cheaper than competitors and is a US company hosted on US based hyperscaler clouds this is kind of a no brainer. Hard for most businesses to justify otherwise.
[0] I am constantly surprised how much work pay-as-you-go with DeepSeek / MiMo will get done. I've barely crossed $2 each in a month of use (~200m tokens).
Re: Advancing the price-performance frontier with GPT‑5.6
#179> In a compute-constrained world where model demand is growing faster than capacity I don't buy it. There have been recent weeks where some of the mid-level models (Hy3, Laguna M.1) are free (true for parts of June and July, see Hy3 in Cyan) . Even then the total token usage appears to be reaching a steady-state. https://openrouter.ai/rankings#top-models ^ the first graph is tokens per week across all models I guess…
My company checks the models and pays for Opus through AWS.
You still send the WHOLE context of whatever you want to do to a random endpoint on the internet. If you want to write a good email, you give that context your email address, names, the reason for it etc.
Big companies don't randomly use some random api endpoint to do so.
Anthropics quarerly revenue is still growing very fast. I don't think we have seen even the real potenzial of it yet at all.
Not only are still a lot of countries missing which do not even use anthropic or any other frontier model yet but also all the agentic based solutions enterprise companies are currently building on mass (at least in my industry)