Earlier quoted context omitted.
Rubin has 22TB/s of memory bandwidth vs Blackwell's 8TB/s. NVLink 6 doubles interconnect speed. Plus they're moving to 3nm from ~4nm. (Previously this comment said Rubin did native NVFP4, but Blackwell does too! Rubin just also trains with native NVFP4, which Blackwell does not.)
Blackwell supports nvfp4 natively.
Performance per dollar is getting faster and cheaper
31–40 of 151 posts
Re: Performance per dollar is getting faster and cheaper
#32Re: Performance per dollar is getting faster and cheaper
#33I think we should make it illegal to not specify the quantization in the headline for these types of posts.
Re: Performance per dollar is getting faster and cheaper
#34Re: Performance per dollar is getting faster and cheaper
#35Re: Performance per dollar is getting faster and cheaper
#36I'm not surprised to see competition with Blackwell. Rubin is 5x faster than Blackwell at inference - Blackwell is the last generation Nvidia didn't optimize specifically for inference. If I'm missing something, please let me know!
Re: Performance per dollar is getting faster and cheaper
#37Can you folks add performance per watt as a metric to these comparisons, I honestly want to understand where AMD fits in the stack in terms of actual performance to dollars. I have had talks with companies wanting to build data centers outside of US and find it hard to source anything Nvidia in sufficient capacity and scale. If AMD is competitive performance per watt and roughly reliable in terms of software support…
> I have never seen a company use AMD outside of wafer and a couple others mostly in US. Just because you haven't seen it doesn't mean it doesn't exist. We've serviced over 700 customers on our MI300x.
Re: Performance per dollar is getting faster and cheaper
#38It's genuinely neat that AI can find the right optimization pathways in an AMD inference server to unlock this but at the same token (pun-intended) this is a classic case of benchmark hacking that doesn't stand up to real-world application.
Re: Performance per dollar is getting faster and cheaper
#39This is very interesting and yet not at the same time. This looks to be optimized for single-stream LLM traffic which is not viable to serve in a production setting. It's only interesting to hobbyists that want to run the model locally. It's genuinely neat that AI can find the right optimization pathways in an AMD inference server to unlock this but at the same token (pun-intended) this is a classic case of benchmark…
Re: Performance per dollar is getting faster and cheaper
#40This is very interesting and yet not at the same time. This looks to be optimized for single-stream LLM traffic which is not viable to serve in a production setting. It's only interesting to hobbyists that want to run the model locally. It's genuinely neat that AI can find the right optimization pathways in an AMD inference server to unlock this but at the same token (pun-intended) this is a classic case of benchmark…