Live data from Hacker News

Who uses Google TPUs for inference in production?

news.ycombinator.com

11–20 of 50 posts

Re: Who uses Google TPUs for inference in production?

#11
post #9

Google is using them in prod. I think they're so hungry for chips internally that cloud isn't getting much support in selling them.

I would guess that Google's vertexAI managed solution uses TPUs. Also Google uses them internally to train and infer for all their research products.

80 to 90% are consumed internally !! Only from version 5 it is planned to be customer focussed !!

Re: Who uses Google TPUs for inference in production?

#13
post #9

Google is using them in prod. I think they're so hungry for chips internally that cloud isn't getting much support in selling them.

I would guess that Google's vertexAI managed solution uses TPUs. Also Google uses them internally to train and infer for all their research products.

While you can use TPUs with vertexai, it's just virtual machines - you can have one with an nvidia card if you like.

Re: Who uses Google TPUs for inference in production?

#17

Google is using them in prod. I think they're so hungry for chips internally that cloud isn't getting much support in selling them.

They're getting swallowed up by Anthropic and the other huge spenders:

https://www.prnewswire.com/news-releases/google-announces-ex...

"Partnership includes important new collaborations on AI safety standards, committing to the highest standards of AI security, and use of TPU v5e accelerators for AI inference "

Re: Who uses Google TPUs for inference in production?

#18
post #6

We've previously tried and almost always regretted the decision. I think the tech stack needs another 12-18 months to mature (doesn't help that almost all work ex Google is being done in torch).

> I think the tech stack needs another 12-18 months to mature

Google has been doing AI before any other company even thought about it. They are on the 6th generation of TPU hardware.

I don't think there is any maturity issue, just an availability issue because they are all being used internally.

Re: Who uses Google TPUs for inference in production?

#19
post #18
post #6

We've previously tried and almost always regretted the decision. I think the tech stack needs another 12-18 months to mature (doesn't help that almost all work ex Google is being done in torch).

> I think the tech stack needs another 12-18 months to mature Google has been doing AI before any other company even thought about it. They are on the 6th generation of TPU hardware. I don't think there is any maturity issue, just an availability issue because they are all being used internally.

100% agree, if I have access to the TPU team internally, it'll be very easy to use in production.

If you aren't internal, the documentation, support, and even just general bug fixing is impossible to get.

Post reply on HN