Exploring inference memory saturation effect: H100 vs. MI300x
1–10 of 13 posts
Re: Exploring inference memory saturation effect: H100 vs. MI300x
#2Re: Exploring inference memory saturation effect: H100 vs. MI300x
#3It would be interesting to know the outright prices for those systems as well as their hourly rental rates at the moment.
Re: Exploring inference memory saturation effect: H100 vs. MI300x
#4Bit weird to show $ per 1M tokens and not include the actual costs of the systems anywhere. It would be interesting to know the outright prices for those systems as well as their hourly rental rates at the moment.
Re: Exploring inference memory saturation effect: H100 vs. MI300x
#5Bit weird to show $ per 1M tokens and not include the actual costs of the systems anywhere. It would be interesting to know the outright prices for those systems as well as their hourly rental rates at the moment.
Hot Aisle: https://hotaisle.xyz/pricing/ Lambda: https://lambdalabs.com/service/gpu-cloud#pricing
Re: Exploring inference memory saturation effect: H100 vs. MI300x
#6Is there some intuition as to why 3.1 might outperform 3.2 on MI300X?
Re: Exploring inference memory saturation effect: H100 vs. MI300x
#7Great read. Have you compared performance with other Llama models (3, 3.2) or have you just done benchmarking with 3.1? Is there some intuition as to why 3.1 might outperform 3.2 on MI300X?
Re: Exploring inference memory saturation effect: H100 vs. MI300x
#8May need a neural net with a third dimension, namely 3D.
Let's wait for the outcome, but they should have a look at this regime in the human brain.
Re: Exploring inference memory saturation effect: H100 vs. MI300x
#9Re: Exploring inference memory saturation effect: H100 vs. MI300x
#10Glad to see hot aisle again providing support to the AMD research community!