Live data from Hacker News

Performance per dollar is getting faster and cheaper

wafer.ai

51–60 of 151 posts

Re: Performance per dollar is getting faster and cheaper

#51

I'm not surprised to see competition with Blackwell. Rubin is 5x faster than Blackwell at inference - Blackwell is the last generation Nvidia didn't optimize specifically for inference. If I'm missing something, please let me know!

It's very unclear what's special in Rubin to be optimized for inference? I can see disaggregated bit (with having separate prefill and decoding nodes), but what else?

Lot more SMs & Tensor Cores for NVFP4 going by the looks of it.

Re: Performance per dollar is getting faster and cheaper

#53

While cool, quantization to FP4 is practically never lossless in actual use. A lot of providers are advertising high TPS on Kimi and GLM, but the models are functionally lobotomized and no longer close to frontier quality. Would love to see this not be true.

from memory, it is like 96-98% of the accuracy.

Re: Performance per dollar is getting faster and cheaper

#54

No word on what this actually means as a consumer. What's the price. Is it lower than NVIDIA serving?

They seem to be serving it at 3x the price while also struggling with maintaining uptime on openrouter; while the vercel router advertizes even bigger speeds but has no clear uptime stats

I guess you really do have to try it at least for some time to actually know

Re: Performance per dollar is getting faster and cheaper

#55
post #53

While cool, quantization to FP4 is practically never lossless in actual use. A lot of providers are advertising high TPS on Kimi and GLM, but the models are functionally lobotomized and no longer close to frontier quality. Would love to see this not be true.

from memory, it is like 96-98% of the accuracy.

Accuracy isn't a meaningful metric here without reference to a specific task.

Re: Performance per dollar is getting faster and cheaper

#56
post #2

Can you folks add performance per watt as a metric to these comparisons, I honestly want to understand where AMD fits in the stack in terms of actual performance to dollars. I have had talks with companies wanting to build data centers outside of US and find it hard to source anything Nvidia in sufficient capacity and scale. If AMD is competitive performance per watt and roughly reliable in terms of software support…

Typically any company that can’t get Nvidia to fill their orders will have at least some AMD.

Re: Performance per dollar is getting faster and cheaper

#57

Earlier quoted context omitted.

hi yes it’s not optimized for single stream it’s optimized for total node throughput

Oh, that's much better then. A good metric to share is the tokens per second per user for the node rather than the total throughput of the node. It disambiguates what's being optimized for much better than your blog post currently does.

sounds good feedback taken, thanks beffjezos

Re: Performance per dollar is getting faster and cheaper

#59
post #48

Earlier quoted context omitted.

A DGX B200 costs like ~$0.5 M and uses around 14 kW. If you plan to run it straight for 8 years 100% max usage thats around 1 GWhr. A gigawatt hour is a lot of energy but its not that much compared to the price of the actual machine. In Germany for example with its expensive energy thats about €100k worth, which spread over 8 years is pretty minor compared to the up front half mill. The real issue with high power con…

It’s more than power supply. Cooling and ventilation becomes a MUCH bigger deal at rack scale, and that costs electricity too.

Cooling demand is only fractional with respect to the load: cooling 1MW of heat will only cost a few 10's to low 100's of kW, depending on the specifics. 10-20% overhead on cooling is probably a close enough estimate for napkin math.

Re: Performance per dollar is getting faster and cheaper

#60
post #53

While cool, quantization to FP4 is practically never lossless in actual use. A lot of providers are advertising high TPS on Kimi and GLM, but the models are functionally lobotomized and no longer close to frontier quality. Would love to see this not be true.

from memory, it is like 96-98% of the accuracy.

And that 2%-4% makes all the difference.
Post reply on HN