Live data from Hacker News

Exploring inference memory saturation effect: H100 vs. MI300x

dstack.ai

1–10 of 13 posts

Re: Exploring inference memory saturation effect: H100 vs. MI300x

#4
post #3

Bit weird to show $ per 1M tokens and not include the actual costs of the systems anywhere. It would be interesting to know the outright prices for those systems as well as their hourly rental rates at the moment.

Hot Aisle: https://hotaisle.xyz/pricing/

Lambda: https://lambdalabs.com/service/gpu-cloud#pricing

Re: Exploring inference memory saturation effect: H100 vs. MI300x

#5
post #4
post #3

Bit weird to show $ per 1M tokens and not include the actual costs of the systems anywhere. It would be interesting to know the outright prices for those systems as well as their hourly rental rates at the moment.

Hot Aisle: https://hotaisle.xyz/pricing/ Lambda: https://lambdalabs.com/service/gpu-cloud#pricing

Yes, we should have included the price too—thanks for pointing that out. I forgot about it. Appreciate you adding the links here! And yes, Lambda's and Hot Aisle's prices were used in the calculation.

Re: Exploring inference memory saturation effect: H100 vs. MI300x

#7

Great read. Have you compared performance with other Llama models (3, 3.2) or have you just done benchmarking with 3.1? Is there some intuition as to why 3.1 might outperform 3.2 on MI300X?

In this one we were only using 3.1 405B FP8. We took one model to simplify the setup and were mostly looking at the memory saturation effect. So basically we compared inference metrics of the same model. I suppose comparing 3.1 and 3.2 will be difficult as they are different models entirely. But open to ideas
Post reply on HN