> You can run DOOM Eternal, building the Stadia that Google couldn’t pull off, because the L40S hasn’t forgotten that it’s a graphics GPU. Savage. I wonder if we’ll see a resurgence of cloud game streaming
We're Cutting L40S Prices in Half
21–30 of 31 posts
Re: We're Cutting L40S Prices in Half
#22Prices lowered to $1.25/hr... still 2X vast.ai prices.
Re: We're Cutting L40S Prices in Half
#23Prices lowered to $1.25/hr... still 2X vast.ai prices.
I don't know what platform vast.ai uses but what I have noticed is cpu compute is pretty slow in those. Specifically the tokenization stage was unusually slow for no apparent reason. Had to give that up and use Google cloud for my research project
They run on literally anything someone installs their agent on.
Re: We're Cutting L40S Prices in Half
#24Earlier quoted context omitted.
There are definitely GPU providers where you can buy cheaper L40S hours than us. I'm not entirely sure what their system architectures are, or whether they're just buying in absolutely spectacular volume, because we are cutting pretty close to the bone with our pricing. One cost factor we have that other providers might not have (I'd love to know): we have to dedicate individual racked physical hosts to each group of…
Hadn't heard of vast.ai before and looked into it. The prices seem really good. Then saw "Our software allows anyone to easily become a host by renting out their hardware." Ya, that's a no from me.
Re: We're Cutting L40S Prices in Half
#25Earlier quoted context omitted.
Hadn't heard of vast.ai before and looked into it. The prices seem really good. Then saw "Our software allows anyone to easily become a host by renting out their hardware." Ya, that's a no from me.
It is a little squicky. Kind of love the idea, though, even if I'd be worried about using rando compute someone installed an agent on.
Re: We're Cutting L40S Prices in Half
#26L40S has 48GB of RAM, curious how they're able to run Llama 3.1 70B on it. The weights alone would exceed this. Maybe they mean quantized/fp8? I just had to implement GPU clustering in my inference stack to support Llama 3.1 70b, and even then I needed 2xA100 80GB SXMs. I was initially running my inference servers on fly.io because they were so easy to get started with. But I eventually moved elsewhere because the pr…
Our standard A100 SXM 80GB price is $3.50/hr, for what it's worth.
H100 will also be much faster, especially if you are willing to use fp8. Maybe 3-4x
Re: We're Cutting L40S Prices in Half
#27nice business to be in I guess.
Re: We're Cutting L40S Prices in Half
#28> You can run DOOM Eternal, building the Stadia that Google couldn’t pull off, because the L40S hasn’t forgotten that it’s a graphics GPU. Savage. I wonder if we’ll see a resurgence of cloud game streaming
GeForce Now is pretty great.
Re: We're Cutting L40S Prices in Half
#29Earlier quoted context omitted.
GeForce Now is pretty great.
Is it responsive enough to play action games? Noticeable lag is what kept me off of Stadia.
It is a surprisingly fluid experience.
Re: We're Cutting L40S Prices in Half
#30they buy them at 12 K, so they pay them off in 1 year approx nice business to be in I guess.