Live data from Hacker News

Who uses Google TPUs for inference in production?

news.ycombinator.com

21–30 of 50 posts

Re: Who uses Google TPUs for inference in production?

#22

Google is using them in prod. I think they're so hungry for chips internally that cloud isn't getting much support in selling them.

I think this is right, in part because I've been told exactly this from people who work for Google and their job is to sell me cloud stuff- i.e., they say they have so much internal demand they aren't pushing TPUs for external use. Hence external pricing and support just isn't that great right now. But presumably when capacity catches up they'll start pushing TPUs again.

Re: Who uses Google TPUs for inference in production?

#24
post #4

Earlier quoted context omitted.

Or maybe they are just using nVidia. Who knows ...

Beyond the fact that this is hardly a secret, there’s lots of other signs. 1. They have bought far less from NVidia than other hyper scalers, and they literally can’t vomit without saying “AI”. They have to be running those models on something. They have purchased huge amounts of chips from fabs, and what else would that be? 2. They have said they use them. Should be pretty obvious here. 3. They maintain a whole soft…

They have announced publicly using TPUs for inference, as far back as 2016. They did not offer TPUs for Cloud customers until 2017. The development is clearly driven by internal use cases. One of the features they publicly disclosed as TPU-based was Smart Reply and that launched in 2015. So their internal use of TPUs for inference goes back nearly a decade.

Re: Who uses Google TPUs for inference in production?

#26
post #8

Apparently Midjourney uses it. GCP put out a press release a while ago: https://www.prnewswire.com/news-releases/midjourney-selects-...

The quote from the linked press release is that they do training on TPUv4, while inference is running on GPUs. I have also heard this separately from people associated with Midjourney recently, and that they solely do training on TPUs.

Re: Who uses Google TPUs for inference in production?

#27
post #6

We've previously tried and almost always regretted the decision. I think the tech stack needs another 12-18 months to mature (doesn't help that almost all work ex Google is being done in torch).

I feel like I have been hearing that since V1 TPU. I think Google is the perfect solution because they are teams whose job is to take a model and TPUify it. Elsewhere there is no team, so it's no fun.

Re: Who uses Google TPUs for inference in production?

#28
post #22

Google is using them in prod. I think they're so hungry for chips internally that cloud isn't getting much support in selling them.

I think this is right, in part because I've been told exactly this from people who work for Google and their job is to sell me cloud stuff- i.e., they say they have so much internal demand they aren't pushing TPUs for external use. Hence external pricing and support just isn't that great right now. But presumably when capacity catches up they'll start pushing TPUs again.

Feels like a bad point in the curve to try and sell them. “Oh our internal hypecycle is done… we’ll put them in the market now that they’re all worn out.

Re: Who uses Google TPUs for inference in production?

#29
I've seen people connecting these to Raspberry Pis to run local LLMs but I'm not sure how effective it is. Check YouTube for some videos about it.

Speaking of SBCs, prior to the Raspberry Pi, I was looking at the Orange Pi 5 which has a Rockchip RK3588S with an NPU (Neural Processing Unit). This was the first I had heard of such a thing but I was curious how/what exactly it does. Unfortunately, there's very little support for Orange Pi & not a large community for it so I couldn't find any feedback on how well it worked or what it did.

http://www.orangepi.org/html/hardWare/computerAndMicrocontro...

Post reply on HN