I'm not surprised to see competition with Blackwell. Rubin is 5x faster than Blackwell at inference - Blackwell is the last generation Nvidia didn't optimize specifically for inference. If I'm missing something, please let me know!
It's very unclear what's special in Rubin to be optimized for inference? I can see disaggregated bit (with having separate prefill and decoding nodes), but what else?
Performance per dollar is getting faster and cheaper
51–60 of 151 posts
Re: Performance per dollar is getting faster and cheaper
#52Re: Performance per dollar is getting faster and cheaper
#53While cool, quantization to FP4 is practically never lossless in actual use. A lot of providers are advertising high TPS on Kimi and GLM, but the models are functionally lobotomized and no longer close to frontier quality. Would love to see this not be true.
Re: Performance per dollar is getting faster and cheaper
#54No word on what this actually means as a consumer. What's the price. Is it lower than NVIDIA serving?
I guess you really do have to try it at least for some time to actually know
Re: Performance per dollar is getting faster and cheaper
#55While cool, quantization to FP4 is practically never lossless in actual use. A lot of providers are advertising high TPS on Kimi and GLM, but the models are functionally lobotomized and no longer close to frontier quality. Would love to see this not be true.
from memory, it is like 96-98% of the accuracy.
Re: Performance per dollar is getting faster and cheaper
#56Can you folks add performance per watt as a metric to these comparisons, I honestly want to understand where AMD fits in the stack in terms of actual performance to dollars. I have had talks with companies wanting to build data centers outside of US and find it hard to source anything Nvidia in sufficient capacity and scale. If AMD is competitive performance per watt and roughly reliable in terms of software support…
Re: Performance per dollar is getting faster and cheaper
#57Earlier quoted context omitted.
hi yes it’s not optimized for single stream it’s optimized for total node throughput
Oh, that's much better then. A good metric to share is the tokens per second per user for the node rather than the total throughput of the node. It disambiguates what's being optimized for much better than your blog post currently does.
Re: Performance per dollar is getting faster and cheaper
#58even having something like opus 4.8 locally would completely change the landscape
Re: Performance per dollar is getting faster and cheaper
#59Earlier quoted context omitted.
A DGX B200 costs like ~$0.5 M and uses around 14 kW. If you plan to run it straight for 8 years 100% max usage thats around 1 GWhr. A gigawatt hour is a lot of energy but its not that much compared to the price of the actual machine. In Germany for example with its expensive energy thats about €100k worth, which spread over 8 years is pretty minor compared to the up front half mill. The real issue with high power con…
It’s more than power supply. Cooling and ventilation becomes a MUCH bigger deal at rack scale, and that costs electricity too.
Re: Performance per dollar is getting faster and cheaper
#60While cool, quantization to FP4 is practically never lossless in actual use. A lot of providers are advertising high TPS on Kimi and GLM, but the models are functionally lobotomized and no longer close to frontier quality. Would love to see this not be true.
from memory, it is like 96-98% of the accuracy.