Live data from Hacker News

We're Cutting L40S Prices in Half

fly.io

1–10 of 31 posts

Re: We're Cutting L40S Prices in Half

#4

> You can run Llama 3.1 70B — the big Llama — for LLM jobs. That's the medium Llama. Does anyone know if an L40S would run the 405B version?

Hi, I'm the person that wrote that sizing comment in the draft for this article. I have been trying for a while and have been unsuccessful at getting 405B running on any of the GPU machines. I suspect I'd need a raw 8xA100 node to do it at Q4. I doubt there is any reasonable combination of L40s cards that can do it on fly.io. It's just too big. I suspect that in time the 70b model will be brought up to be roughly equivalent, but realistically it's already on the GPT-4 threshold as is. I've found that 70b is more than sufficient in practice.

Re: We're Cutting L40S Prices in Half

#5
post #4

> You can run Llama 3.1 70B — the big Llama — for LLM jobs. That's the medium Llama. Does anyone know if an L40S would run the 405B version?

Hi, I'm the person that wrote that sizing comment in the draft for this article. I have been trying for a while and have been unsuccessful at getting 405B running on any of the GPU machines. I suspect I'd need a raw 8xA100 node to do it at Q4. I doubt there is any reasonable combination of L40s cards that can do it on fly.io. It's just too big. I suspect that in time the 70b model will be brought up to be roughly equ…

Be that as it may, Llama 3.1 70B is not the big Llama.

Re: We're Cutting L40S Prices in Half

#6
> You can run DOOM Eternal, building the Stadia that Google couldn’t pull off, because the L40S hasn’t forgotten that it’s a graphics GPU.

Savage.

I wonder if we’ll see a resurgence of cloud game streaming

Re: We're Cutting L40S Prices in Half

#7
post #5
post #4

Earlier quoted context omitted.

Hi, I'm the person that wrote that sizing comment in the draft for this article. I have been trying for a while and have been unsuccessful at getting 405B running on any of the GPU machines. I suspect I'd need a raw 8xA100 node to do it at Q4. I doubt there is any reasonable combination of L40s cards that can do it on fly.io. It's just too big. I suspect that in time the 70b model will be brought up to be roughly equ…

Be that as it may, Llama 3.1 70B is not the big Llama.

I fixed it.

Re: We're Cutting L40S Prices in Half

#8
post #2

Prices lowered to $1.25/hr... still 2X vast.ai prices.

There are definitely GPU providers where you can buy cheaper L40S hours than us. I'm not entirely sure what their system architectures are, or whether they're just buying in absolutely spectacular volume, because we are cutting pretty close to the bone with our pricing.

One cost factor we have that other providers might not have (I'd love to know): we have to dedicate individual racked physical hosts to each group of GPUs we deploy, because we don't (/can't, depending on how you think about systems security) allow GPU-enabled workloads to share hardware with non-GPU-enabled workloads, and we don't allow anyone to share kernels.

But like we said in the post: we're still figuring this stuff out. What we know is: at the same price level, we're consistently sold out of A10 inventory.

Re: We're Cutting L40S Prices in Half

#10
post #6

> You can run DOOM Eternal, building the Stadia that Google couldn’t pull off, because the L40S hasn’t forgotten that it’s a graphics GPU. Savage. I wonder if we’ll see a resurgence of cloud game streaming

Is the services that PlayStation Now uses publicly known? That's the only streaming service I've used so far.
Post reply on HN