I am super interested in AI on a personal level and have been involved for a number of years. I have never seen a GPU crunch quite like it is right now. To anyone who is interested in hobbyist ML, I highly highly recommend using vast.ai
Additional clouds:
For H100s and A100s - lambda, fluidstack, runpod. Also coreweave and crusoe and oblivus and latitude
For non a/h100s: vast, Tensordock, also runpod here too
Nat Friedman and Daniel Gross setup a 2,512 H100 cluster [1] for their startups, with a very similar “shared” model. Might be interesting to connect with them.
Realism doesn't work in business. Business success requires 10 people to try for 1 person to succeed. If those 10 people were realists, they wouldn't try.
Depends how much that one person wins and how much the others lose.
I know AWS/GCP/Azure have overhead and I understand why so many companies choose to go bare metal on their ops. I personally rarely think it's worth the time and effort, but I get that with scale saving can be substantial.
But for AI training? If the public cloud isn't competitive even for bursty AI training, their margins are much higher than I anticipated.
OP mentions 10-20x cost reduction? Compared to what? AWS?
Yeah, it’s a silly branding thing. One TPU (not even a pod, just a regular old TPUv2) has 96 CPU cores with 1.4TB of RAM, and that’s not even counting their hardware acceleration. I’d love to buy one.
Huh, this doesn't seem right. Based on #s you seem to be referring to pods but even then I'm not familiar with such a configuration existing. A single TPUv2 chip has 1 core and 8gb of memory. A single device comes in the v2-8 configuration with 8 cores and 64gb of memory. Pod variants come in v2-32 to v2-512 configurations.
A single TPUv2 host has 8 TPU cores with 64GB of total HBM (8GB per core), but like GPUs, TPUs can't directly access a network, so the host also needs CPUs and standard RAM to send data to them. They are fast, and the host has to be fast enough to keep them fed with data, so the host is pretty beefy. But FWIW, a TPUv2 host has somewhere around 330GB of RAM, not 1.4TB.
TPU pod is not sold by google, edge tpu is different
So the cloud TPUs are more powerful...? Or what are you saying?
Edge TPUs are low cost, low power inference devices the size of a dime. I have a hundred of them sitting in a closet. (Alas. Anyone want to buy 100 coral minis? :-)
The TPUs you rent that are being discussed here are capable of training, consume hundreds of watts and have a heatsink bigger than your fist and really spectacular network links. They're analogous to Nvidia's highest end GPUs from a "what can you do with them" perspective.
Both are custom chips for deep learning but they're completely different beasts.
So the cloud TPUs are more powerful...? Or what are you saying?
Edge TPUs are low cost, low power inference devices the size of a dime. I have a hundred of them sitting in a closet. (Alas. Anyone want to buy 100 coral minis? :-) The TPUs you rent that are being discussed here are capable of training, consume hundreds of watts and have a heatsink bigger than your fist and really spectacular network links. They're analogous to Nvidia's highest end GPUs from a "what can you do with…
Can I hook a microphone up to a Coral Mini and run Whisper? I'd love to have a home assistant that wasn't on the cloud.
As for the rest of them, list them on Amazon and let them do the fulfillment. That $10k of hardware isn't going to sell itself from your closet. (Yet. LLMs are making great strides.)
> You don’t need to take my word for it. Here’s some unfiltered DMs on the subject: https://imgur.com/a/6vqvzXs > Notice how their optimism dries up, and not because I was telling them how bad TRC has become. It’s because their TPUs kept dying. Unless I'm misreading this they sound pretty happy and you sound pessimistic? Their last substantial comment was "I'm sure Zak could hook you up with something better"?
TRC is supposed to be the “something better”. This insider TPU stuff is for the birds. If TRC can only offer 4 hours with no preemptions, that’s fine, but they need to be up front about that. Saying that TPUs preempt every 24 hours and then killing them off after 45 minutes is… not very productive. As for their comments, the third screenshot is the key; they’re agreeing that the situation is bad. They’re a friend, an…
I don't have a qualified opinion on the subject of TPU availability.
I'm just pointing out that your summary of the DMs ("Notice how their optimism dries up, and not because I was telling them how bad TRC has become. It’s because their TPUs kept dying") is the opposite of what the DMs show.