Earlier quoted context omitted.
imagine writing that on your resume > reduced inference cost by 20 percent saving company x billion dollars per month
Contributed to. Can't be some IC who made a few nice PRs
Advancing the price-performance frontier with GPT‑5.6
221–230 of 424 posts
Re: Advancing the price-performance frontier with GPT‑5.6
#222This feels like the dialup->broadband transition to me. I was already a huge proponent of Luna for things like deep research. Being able to run 5x more for the same cost is simply bananas. We are already running 10 parallel agents for hypothesis generation. I cannot imagine 50. The statistics become much more interesting & powerful when you can run so many samples of the exact same prompt+model without breaking the b…
How do you run 'deep research'?
Re: Advancing the price-performance frontier with GPT‑5.6
#223> Starting today, GPT‑5.6 Luna, our fastest and most affordable model, will cost 80% less, I don't have the words. I genuinely thought we were in a stage where we were plateauing and going in for 5-10% improvements over months. Seeing spikes like this makes me question about where the floor really is.
Re: Advancing the price-performance frontier with GPT‑5.6
#22480% price cut for luna is a very aggressive pricing move makes it by far the best choice for most workloads that do not need bleeding edge intelligence (reminder: luna can be comparable to opus 5!)
Re: Advancing the price-performance frontier with GPT‑5.6
#225> Starting today, GPT‑5.6 Luna, our fastest and most affordable model, will cost 80% less, I don't have the words. I genuinely thought we were in a stage where we were plateauing and going in for 5-10% improvements over months. Seeing spikes like this makes me question about where the floor really is.
> Seeing spikes like this makes me question about where the floor really is. You mean they increased the price and then cut it back and now it is amazing? Luna had a price hike vs mini (its previous replacement). The cut now just puts it back in that ball park. Not that this isn't good news, but what's impressive?
Re: Advancing the price-performance frontier with GPT‑5.6
#226To fix this you currently need to make your own copy of the bundled model catalog [2] and opt Luna into MultiAgent V2.
1. https://github.com/openai/codex/issues/32031
2. https://github.com/openai/codex/issues/32031#issuecomment-51...
Re: Advancing the price-performance frontier with GPT‑5.6
#227> In a compute-constrained world where model demand is growing faster than capacity I don't buy it. There have been recent weeks where some of the mid-level models (Hy3, Laguna M.1) are free (true for parts of June and July, see Hy3 in Cyan) . Even then the total token usage appears to be reaching a steady-state. https://openrouter.ai/rankings#top-models ^ the first graph is tokens per week across all models I guess…
Stop rejecting what we have been collectively telling you guys! LLM providers are profitable, and have been for awhile!
Re: Advancing the price-performance frontier with GPT‑5.6
#228Earlier quoted context omitted.
To be fair we don't really know in terms of prices what's real and what's just investor subsidised attempts at market capture at this point. It could well be OpenAI's attempt to drown Anthropic while they've got the halo product if they feel they've got deeper pockets.
Enterprises implemented spending caps and inference providers are lowering prices. Seems they are jockeying for market share.
Re: Advancing the price-performance frontier with GPT‑5.6
#229Seems like they're working to destroy the local LLM argument. Right now Haiku is $1/$5 in/out. You can grind out $12,000 worth of haiku (or arguably, sonnet) class tokens in about 5 months on a Blackwell RTX 6000 96GB especially if using concurrency. BUT, but, if you use a g6e.xlarge on aws it's now more expensive than buying tokens from OpenAI @ $0.20/$1.20. It also destroys "the Mac Mini argument", pushing the ROI…
The local LLM argument never really held water tbh. You can get surprisingly good performance for lightweight tasks locally, but you're just fighting economies of scale if you're going trying to beat a datacenter on cost.