Live data from Hacker News

Performance per dollar is getting faster and cheaper

wafer.ai

31–40 of 151 posts

Re: Performance per dollar is getting faster and cheaper

#31

Earlier quoted context omitted.

Rubin has 22TB/s of memory bandwidth vs Blackwell's 8TB/s. NVLink 6 doubles interconnect speed. Plus they're moving to 3nm from ~4nm. (Previously this comment said Rubin did native NVFP4, but Blackwell does too! Rubin just also trains with native NVFP4, which Blackwell does not.)

Blackwell supports nvfp4 natively.

You're right - Rubin is better at NVFP4 training, not inference, thank you for catching me!

Re: Performance per dollar is getting faster and cheaper

#34
post #3

Do these providers have 80+% gross margins or is something eating into them? Maybe utilization?

hi i work at wafer. no the margins are lower averaging at about ~40%. utilization is one of the highest order bits in determining margins here, yes.

[dead]

Re: Performance per dollar is getting faster and cheaper

#35

Earlier quoted context omitted.

Blackwell supports nvfp4 natively.

You're right - Rubin is better at NVFP4 training , not inference, thank you for catching me!

What does it mean it's better at nvfp4 training? What's different between training and inference to make this true?

Re: Performance per dollar is getting faster and cheaper

#36

I'm not surprised to see competition with Blackwell. Rubin is 5x faster than Blackwell at inference - Blackwell is the last generation Nvidia didn't optimize specifically for inference. If I'm missing something, please let me know!

It's very unclear what's special in Rubin to be optimized for inference? I can see disaggregated bit (with having separate prefill and decoding nodes), but what else?

Re: Performance per dollar is getting faster and cheaper

#37
post #2

Can you folks add performance per watt as a metric to these comparisons, I honestly want to understand where AMD fits in the stack in terms of actual performance to dollars. I have had talks with companies wanting to build data centers outside of US and find it hard to source anything Nvidia in sufficient capacity and scale. If AMD is competitive performance per watt and roughly reliable in terms of software support…

> I have never seen a company use AMD outside of wafer and a couple others mostly in US. Just because you haven't seen it doesn't mean it doesn't exist. We've serviced over 700 customers on our MI300x.

[deleted]

Re: Performance per dollar is getting faster and cheaper

#38
This is very interesting and yet not at the same time. This looks to be optimized for single-stream LLM traffic which is not viable to serve in a production setting. It's only interesting to hobbyists that want to run the model locally.

It's genuinely neat that AI can find the right optimization pathways in an AMD inference server to unlock this but at the same token (pun-intended) this is a classic case of benchmark hacking that doesn't stand up to real-world application.

Re: Performance per dollar is getting faster and cheaper

#39

This is very interesting and yet not at the same time. This looks to be optimized for single-stream LLM traffic which is not viable to serve in a production setting. It's only interesting to hobbyists that want to run the model locally. It's genuinely neat that AI can find the right optimization pathways in an AMD inference server to unlock this but at the same token (pun-intended) this is a classic case of benchmark…

You got it backwards; it's ~200 on single stream so the 2,600 is achieved with ~13 streams.

Re: Performance per dollar is getting faster and cheaper

#40

This is very interesting and yet not at the same time. This looks to be optimized for single-stream LLM traffic which is not viable to serve in a production setting. It's only interesting to hobbyists that want to run the model locally. It's genuinely neat that AI can find the right optimization pathways in an AMD inference server to unlock this but at the same token (pun-intended) this is a classic case of benchmark…

hi yes it’s not optimized for single stream it’s optimized for total node throughput
Post reply on HN