Live data from Hacker News

Advancing the price-performance frontier with GPT‑5.6

openai.com

301–310 of 424 posts

Re: Advancing the price-performance frontier with GPT‑5.6

#301

Earlier quoted context omitted.

Yep it will be ASICs and DSPs all over again. Orders of magnitude changes.

So which shovels companies are the ones to watch for burnt in silicon models ?

$CBRS - Cerebras Systems

There is other in the space, Groq and Sambanova are both private companies attempting to develop their own technology.

Re: Advancing the price-performance frontier with GPT‑5.6

#302
post #159

Earlier quoted context omitted.

Holy crap, I was not prepared for how fast it responded. I just wrote "Just wanted to see how fast you are! Can you write me a quick story about a tiger who lives inside a block of cheese the size of a house?" I pressed Enter, and the response was instant . > Generated in 0.037s • 14,205 tok/s This is unbelievable.

It's crazy. Are they doing any precomputing as you type, I wonder if you paste a block of text is it the same speed.

The magic is in the fact that they essentially have an ASIC llm device. There is no other trickery. The problem they will face is that it is actually locked in silicon, so upgrading models will be difficult, and likely require new hardware each time.

Re: Advancing the price-performance frontier with GPT‑5.6

#303

Earlier quoted context omitted.

I don't see why it should be all that difficult. All you have to do is first find a library that implements a decent solution to the halting problem and you're off to the races.

You can use my script p-noteq-np.sh too if that helps.

I checked out your script but it looks like it's just a wrapper for ! ( p-eq-np.sh )

Re: Advancing the price-performance frontier with GPT‑5.6

#304
post #3

> The kernel work helped reduce the end-to-end cost of serving the model by 20%, while its experiments increased token-generation efficiency by more than 15%. If the cost of serving GPT-5.6 just dropped by 20%, does that add up to literally billions of dollars in savings per month? We know Anthropic spend $1.25 billion renting inference capacity from SpaceX (in two Colossus datacenters) from the SpaceX IPO, but we do…

Some numbers: https://www.wheresyoured.at/exclusive-openai-financials/ If those numbers are accurate, I don't think 20% is a really, really big deal. It's like saying "we're digging our grave 20% slower." Ok, but they're still digging! Or, different analogy, if I'm going broke because I lost my job due to executive AI psychosis, cancelling my netflix subscription doesn't really change the math of not being able to af…

Those finanicals show OpenAI makes good money on inference.

Re: Advancing the price-performance frontier with GPT‑5.6

#305

"Half the money I spend on advertising is wasted; the trouble is I don't know which half." -John Wanamaker This applies even more strongly to model choosing. I know for a fact that majority of my work doesn't require a very strong model, but separating the trivial and non-trivial tasks is a famously hard problem (if at all decidable).

I don't see why it should be all that difficult. All you have to do is first find a library that implements a decent solution to the halting problem and you're off to the races.

Its funny because you can write a halting problem oracle by calling out to an LLM and have it return yes / no / not sure and get it to work reliably for almost all real code, like that is an entirely practical thing to do in 2026.

All we need now is some sort of program to evaluate halting problem oracles...

Re: Advancing the price-performance frontier with GPT‑5.6

#306

Earlier quoted context omitted.

https://taalas.com/ has done it already for a wildly obsolete model. 14000 tokens per second. https://chatjimmy.ai/ is their interactive. Tiny context, very dumb, but absurdly fast. Imagine this as a tool call for claude code for trivial changes - the tool call from the harness takes longer than the execution.

Wow. This is absolutely wild. I didn't expect that. If we get to anywhere near this speed for the equivalent of the current models... I don't even know what to think about that future.

what do you mean "if"? of course we will, and the models will be smarter as well

Re: Advancing the price-performance frontier with GPT‑5.6

#307

"Half the money I spend on advertising is wasted; the trouble is I don't know which half." -John Wanamaker This applies even more strongly to model choosing. I know for a fact that majority of my work doesn't require a very strong model, but separating the trivial and non-trivial tasks is a famously hard problem (if at all decidable).

You just need a very strong frontier model to do triage of your tasks. /s

I mean, yeah. Alternatively I’ve seen the approach of starter with the cheaper/faster model but give it a tool to handoff or talk to a strong model if it gets stuck.

Re: Advancing the price-performance frontier with GPT‑5.6

#308
post #77

Earlier quoted context omitted.

Vera Rubin will be hitting racks very soon, and this is purported to have a 10x improvement in token throughput per megawatt. Of course, old chips don't get replaced with new chips overnight, but I don't think we're anywhere near the floor yet.

AMD MI400 series is already shipping to customers (basically everybody) and it is crazy fast (8x to 10x faster than the previous gen and beats published Vera numbers in FP8, loses in FP4) and 432 GB per chip. 72 chip unified rack architecture (Helios) already shipping and projected to also beat Vera in NVL72. MI500 series is supposedly already taping out and they're claiming massive increases (we'll find out end of 2…

Absolutely true although at some point it's not just raw numbers but also the kernels that run matmuls and there seems (from outsider perspective) to have been more optimization in the cuda kernels

Re: Advancing the price-performance frontier with GPT‑5.6

#309
post #292
post #255

Earlier quoted context omitted.

I'm not who you responded to and I don't have any info on Google. Nor can I explain in detail due to NDAs. But multiple major players are working on something along the lines of what the parent is alluding to. The "edge" AI landscape (in particular, what you can do with ~5W) is going to be nuts in about 18 months.

How will this affect the newly build data centers? What effect do you think it will have on memory prices?

My uneducated guess says, not much. For running massive models you still need a ton of high-bandwidth interconnects between many individual chips/GPUs/etc since you need to do math across a few TB worth of weights. That's simply going to require more power (and more die area in I/O, and therefore more cost). Being able to run small models in tiny power envelopes is incredibly useful to people, but I believe it will be covering a different niche than what datacenters can provide. Likewise, you'll still need crazy amounts of high-end memory to populate whatever goes in these datacenters.

The only thing that will crash prices is reduced demand (duh) or, more interestingly, increased production. In particular, if CXMT is able to get their DDR5 fabs up to a reasonably high yield, that could add some downward price pressure (as could government subsidies). As well, if Micron/Kingston/Hynix think that CXMT is going to start cutting into their market share, they might be willing to either increases supply or drop prices. Unfortunately CXMT looks to be taking quite a while to get their new fab up to max capacity so that may take a year+ before anything manifests.

If you're interested in following the (publicly available) info on these sorts of things, check out what companies like Axelera, DeepX, and MemoryX are doing today and have on their roadmaps, as well as the sorts of chips/SoCs Qualcomm, Kinara (now NXP), and Ambarella currently have announced (or have on the market). And remember, that pretty much all of these chips on the market today were in initial development more or less when ChatGPT first launched. If you knew what you knew today (or a year ago) about what requirements current- and next-generation models would have (from a silicon perspective), what might you do differently? Think for instance, host system interconnects, amount and speed of on-package or on-die memory, image/video decode capabilities, int8 vs fp8 vs fp16 vs bf16 compute units, etc. And, consider that most "AI" stuff in development a few years ago was all 15nm or 12nm - because who was gonna pay big money to get fab capacity at 3nm to run some object detection models? So most of the stuff on the market today is on very old nodes and therefore not super power efficient.

Re: Advancing the price-performance frontier with GPT‑5.6

#310
post #14

> Starting today, GPT‑5.6 Luna, our fastest and most affordable model, will cost 80% less, I don't have the words. I genuinely thought we were in a stage where we were plateauing and going in for 5-10% improvements over months. Seeing spikes like this makes me question about where the floor really is.

this type of thing usually means you are the product

These are the API rates. They don’t train on prompts from the API. In what sense are you the product?
Post reply on HN