Live data from Hacker News

Powering cost-efficient AI inference at scale with Cloud TPU v5e on GKE

cloud.google.com

21–28 of 28 posts

Re: Powering cost-efficient AI inference at scale with Cloud TPU v5e on GKE

#22

AI is a big deal this year. Google is one of the biggest tech companies. Yet a post about AI at Google gets... No comments at all. Back in 2005, this would have been the most talked about news of the day. Now, nobody cares what Googles up to. They lost their way.

My take as a GCP employee: Amazon has really dominated the startup world, but google cloud seems to really be hitting its stride with big enterprise tech companies.

Between its support for K8s, networking stack/infra, and TPUs/ML support we have a lot of huge well known tech companies knocking down our door.

But none of those things really matter as a differentiator for the average startup, scale up, or small business. So all in all, i’m not surprised the majority of HN engineers just aren’t excited by GCP.

I have no knowledge of Cloud’s long term strategy, but it wouldn’t surprise me if by leaning into its success in big tech, they are neglecting the startup scene in terms of marketing or building out new features.

Re: Powering cost-efficient AI inference at scale with Cloud TPU v5e on GKE

#23
post #15

AFAIK Google is in the middle of massive infrastructure investment in hardware which is of similar performance to Nvidia, but they will own the whole stack. They will also be able to deploy an order of magnitude more compute than openAI to LLMs in 2024/25. LLM tech also seems to be easily swapped between providers. So my view is that the value proposition is in the hardware designers and manufacturers (Nvidia, tsmc,…

My guess is that google ran the numbers and isn’t willing to lose money like openai is, so they have been slower to develop things as cool as openai, leading to a perception that they can’t do it. Hence the lack of excitement.

But I think that’s unlikely, and once it’s not money-losing, google will be right there. Which, as you say, they might have an edge on due to hardware.

Re: Powering cost-efficient AI inference at scale with Cloud TPU v5e on GKE

#24
post #15

AFAIK Google is in the middle of massive infrastructure investment in hardware which is of similar performance to Nvidia, but they will own the whole stack. They will also be able to deploy an order of magnitude more compute than openAI to LLMs in 2024/25. LLM tech also seems to be easily swapped between providers. So my view is that the value proposition is in the hardware designers and manufacturers (Nvidia, tsmc,…

> Why are more people not excited by what is happening at Google? What am I missing?

I'd love contra-indicators. I'm not plugged in to this community nearly as muc has I am the rest of tech. But the general scuttlebutt I got was that access to TPU4 was quite competitive. It didn't seem to be like a cloud resource where you could just go use it: it took work and effort and time and many other people could butt in line ahead of you all the time to get TPUv4. Apologies for relaying hearsay, but if this is far better tech, it feels like Google is going to have to solve the scale-out concerns exponentially harder if TPUv5e is to actually* have an effect on most practitioners. It's my hope that these negative concerns were overstated; it's my hope that the situation is better. Does anyone have any idea what we could look at to suss how much of a concern this is?

I think you've also casually stumbled into another issue. To quote you:

> but they will own the whole stack.

This feels like the typical slant against Google, and I think often this is a disservice.

I personally attribute much better to Google. To me, Google has been one of the few companies that still has some of the intertwingularity in their ethos, that still has some Cluetrain Manifesto blood in their veins where they grok that the way to be a winning player is to bring a ton of selfless. But we keep seeing spin that makes Google look egocentric & self-biased & self-dealing, in so many ways. Any bad Google anywhere outshadows the rest of the good Google; Google is a massive goliath powered by the vastest ad market on the planet, but where-as most companies constantly focus on mining value of their ecosystem & imho Google is one of the rare few companies that groks that the existential stakes of creating open healthy accessible ecosystems. Imho they have almost never stacked the deck to be overlords (fuck you very very much MV3, you self-dealing app-store owners how casually rescind so many dimensions of freedom so easily for yourselves) but instead espoused a multi-polar world, one made up of many independent empowered factions. If anything, I think much of the scorn laid at Chrome's feet, for example, actually reflects on the rest of the world's inability to be as good, to serve as even remotely adequate patrons to the web: governments everywhere should be funding web development, web specs, web standards. There's just so few players & that's not Google's fault. The bread & butter of open information society is all too moldy & stale in most areas, too few players. Still, I've seen Google persist, try. I find it admirable. I see this in many ways, but the faults & criticisms are always such drastic lightning rods for anger.

No anger here yet, but: there's hardly a better example of Google trying to play well than today's Keras 3.0 announcement. 'Hello, we are Google, we've optimized the ever loving hell out of our existing machine learning platform to target a whole bunch of other environments, make them 2D-grid data & model parallel across a variety of frameworks targeting by far the widest range of hardwares. And so now, if you target Keras, any of these other platforms will work intercompatibly very well.' Once again, Google steps in the ring to carry water for the world, to make practioner's lives much easier. Sure, this enables TPUv5e to be easily targetted, via a variety of different code-styles; that is self interested. But it also enables a degree of portability out of TPU & across systems that seems magical, that is extremely late bound, that permits permits permits. I'm harping on just a small bit you said, apologies, but in contrast to your 'own the whole stack', it feels like Google is the only player trying to make sure no one owns the stack, and it seems like they're doing great at it. TPUv5e happens to be one option these actions bolster. https://keras.io/keras_3/ https://news.ycombinator.com/item?id=38446353

Re: Powering cost-efficient AI inference at scale with Cloud TPU v5e on GKE

#26

Fantastic blog! In your post, you highlighted the use of Cloud TPU v5e with GKE for AI inference. How does this setup maintain high performance while managing costs, especially in high-demand scenarios like real-time data processing or live interactions?

Thank you! Regarding performance, TPU v5e which was benchmarked in the MLPerf results and showed impressive performance/dollar, see https://cloud.google.com/blog/products/compute/performance-p... which has more details. Combined with k8s & GKE, TPU v5e workloads can leverage auto-scaling, for example setting up autoscaling based on traffic, so workloads can scale down rapidly when not in use, increasing cost efficien…

Can I use LLMs to make more convincing sock puppet questions?

Re: Powering cost-efficient AI inference at scale with Cloud TPU v5e on GKE

#27
post #19
post #4

Earlier quoted context omitted.

Google, structurally, sucks at anything but sucking down Ad money. A good LLM allows one to basically bypass the internet for many use cases and drastically reduces ad time. As cool as LLM capabilities are the thing that delights me most about them is I can get high quality (YMMV) answers from the internet without all the godforsaken ad bullshit and wonky websites made only as mazes to keep users in ad gardens. Googl…

LLMs are great at historical data. They however still suck at being kept up to date. I would expect that someone (if not already has) will find a cost efficient way to continuously index new data, but it would be interesting to see how it stacks up with the computing power needed for modern search index pipelines.

Yeah, when I said "many use cases", that mostly means, "use cases that don't require near-term knowledge." Which are the bulk of my internet searches.

But even know, ChatGPT 4 can use search to get update-to-date info and use it's knowledge store to augment the current information.

Re: Powering cost-efficient AI inference at scale with Cloud TPU v5e on GKE

#28
post #4

AI is a big deal this year. Google is one of the biggest tech companies. Yet a post about AI at Google gets... No comments at all. Back in 2005, this would have been the most talked about news of the day. Now, nobody cares what Googles up to. They lost their way.

Google, structurally, sucks at anything but sucking down Ad money. A good LLM allows one to basically bypass the internet for many use cases and drastically reduces ad time. As cool as LLM capabilities are the thing that delights me most about them is I can get high quality (YMMV) answers from the internet without all the godforsaken ad bullshit and wonky websites made only as mazes to keep users in ad gardens. Googl…

"As far as this particular product, it's touting cost efficiency, but that's not really dominating many decisions yet"

Why do these LLM startups have such large investments and such large spending? It's because the infrastructure costs of a GenAI startup is huge. That's most of the money that's changing hands. So, from the VCs down to accelerator platforms cost efficiency is a huge decision factor on all of this.

Post reply on HN