> Starting today, GPT‑5.6 Luna, our fastest and most affordable model, will cost 80% less, I don't have the words. I genuinely thought we were in a stage where we were plateauing and going in for 5-10% improvements over months. Seeing spikes like this makes me question about where the floor really is.
Vera Rubin will be hitting racks very soon, and this is purported to have a 10x improvement in token throughput per megawatt. Of course, old chips don't get replaced with new chips overnight, but I don't think we're anywhere near the floor yet.
Advancing the price-performance frontier with GPT‑5.6
161–170 of 424 posts
Re: Advancing the price-performance frontier with GPT‑5.6
#162Didn't expect that. Luna pricing is crazy now. I don't think there is anything on the market that competes at this price-performance point. For our production app, OpenAI clearly is the best provider now. Their API is very reliable and has many nice features. The price-performance of the model lineup is incredible. We used open weights model via Fireworks for a long time (e.g. Kimi K2.5). Fireworks is a great provide…
OpenAI's APIs are extremely reliable for sure. I don't even remember when the last incident or downtime was.
Re: Advancing the price-performance frontier with GPT‑5.6
#163Earlier quoted context omitted.
Totally possible that humans aren't actually that intelligent.
As well as the existing intelligence being swayed by emotions, hormones, circadian rhythms, stress, peer pressure, propaganda, and survival instincts.
We are not purely rational creatures, thank God. Sometimes those "limiting factors" you listed -- stress, peer pressure, hormones -- are crucial elements of informing the problem solving process and arriving at a decision or a solution that actually works.
All an LLM can do is fulfill a prompt, no matter how misguided, backwards, or incomplete that prompt actually was.
"Go jump off a bridge." Hmm. Dying makes me stressed out. I'm not gonna do that.
Re: Advancing the price-performance frontier with GPT‑5.6
#164> The kernel work helped reduce the end-to-end cost of serving the model by 20%, while its experiments increased token-generation efficiency by more than 15%. If the cost of serving GPT-5.6 just dropped by 20%, does that add up to literally billions of dollars in savings per month? We know Anthropic spend $1.25 billion renting inference capacity from SpaceX (in two Colossus datacenters) from the SpaceX IPO, but we do…
Re: Advancing the price-performance frontier with GPT‑5.6
#165Earlier quoted context omitted.
Vera Rubin will be hitting racks very soon, and this is purported to have a 10x improvement in token throughput per megawatt. Of course, old chips don't get replaced with new chips overnight, but I don't think we're anywhere near the floor yet.
In a data center that is power constrained but not space constrained they could build out new racks and flip the power from the old racks. Wonder if this will lead to moderately used server GPUs on the secondary market someday.
Besides Nvidia Hardware is still sold out and super expensive. Not a single Nvidia consumer GPU got cheaper at all, Nvidia DGX Spark got more expensive too.
It will be swooped of the market the second it hits the market.
Re: Advancing the price-performance frontier with GPT‑5.6
#166Earlier quoted context omitted.
https://taalas.com/ has done it already for a wildly obsolete model. 14000 tokens per second. https://chatjimmy.ai/ is their interactive. Tiny context, very dumb, but absurdly fast. Imagine this as a tool call for claude code for trivial changes - the tool call from the harness takes longer than the execution.
Holy crap, I was not prepared for how fast it responded. I just wrote "Just wanted to see how fast you are! Can you write me a quick story about a tiger who lives inside a block of cheese the size of a house?" I pressed Enter, and the response was instant . > Generated in 0.037s • 14,205 tok/s This is unbelievable.
Re: Advancing the price-performance frontier with GPT‑5.6
#16780% less for Luna is absolutely crazy, in my opinion we may reach a point in the next year where powerful models on the API could potentially be cheaper than subscriptions. Compute just keeps decreasing in price.
Re: Advancing the price-performance frontier with GPT‑5.6
#168- lower input/output token pricing
- the cached token price is $0.0028/Million tokens, which is like 50-90% of tokens
Re: Advancing the price-performance frontier with GPT‑5.6
#169Re: Advancing the price-performance frontier with GPT‑5.6
#170I would pay significantly more to use these models if there was a legal contract that guaranteed they weren't ever terfing them and some way to prove that.