Live data from Hacker News

Google's First Tensor Processing Unit: Architecture

thechipletter.substack.com

181–190 of 197 posts

Re: Google's First Tensor Processing Unit: Architecture

#181
post #29

On the podcast interview now Groq CEO Jonathon Ross did[1] he talked about the creation of the original TPUs (which he built at Google). Apparently originally it was a FPGA he did in his 20% time because he sat near the team who was having inference speed issues. They got it working, then Jeff Dean did the math and the decided to do an ASIC. Now of course Google should spin off the TPU team as a separate company. It'…

The way I see, NVidia only has a few advantages ordered from most important to least: 1. Reserved fab space. 2. Highly integrated software. 3. Hardware architecture that exists today. 4. Customer relationships. but all of these aspects are weak in one way or another: For #1, fab space is tight, and NVidia can strangle its consumer GPU market if it means selling more AI chips at a higher price. This advantage is gone…

I’ve spent the last month deep in GPU driver/compiler world and -

AMD or Apple (Metal) or someone (I haven’t tried Intel’s stuff) just needs to have a single guide to installing a driver and compiler that doesn’t segfault if you look at it wrong, and they would sweep the R&D mindshare.

It is insane how bad CUDA is; it’s even more insane how bad their competitors are.

Re: Google's First Tensor Processing Unit: Architecture

#182

Earlier quoted context omitted.

The way I see, NVidia only has a few advantages ordered from most important to least: 1. Reserved fab space. 2. Highly integrated software. 3. Hardware architecture that exists today. 4. Customer relationships. but all of these aspects are weak in one way or another: For #1, fab space is tight, and NVidia can strangle its consumer GPU market if it means selling more AI chips at a higher price. This advantage is gone…

I’ve spent the last month deep in GPU driver/compiler world and - AMD or Apple (Metal) or someone (I haven’t tried Intel’s stuff) just needs to have a single guide to installing a driver and compiler that doesn’t segfault if you look at it wrong, and they would sweep the R&D mindshare. It is insane how bad CUDA is; it’s even more insane how bad their competitors are.

If you work in hardware and are interested in solving this lemme say this

There are billions of dollars waiting for the first person to get this right. The only reason I haven’t jumped on this myself is a lack of familiarity with drivers.

Re: Google's First Tensor Processing Unit: Architecture

#183

Earlier quoted context omitted.

The answer is far weirder - they had a chat bot, and no one even discussed it in the context of search replacements. They didn’t want to release it because they just didn’t think it should be a product. Only after OpenAI actually disrupted search did they start releasing Gemini/Bard which takes advantage of search.

They were afraid to release it because of unaligned output and hallucinations. ChatGPT showed that people could still get value out of something that wasn’t perfect. E.g. they had this in their labs: https://www.theguardian.com/technology/2022/jun/12/google-en... from July, 2022z

LaMBDA was also briefly available for public testing, but then rapidly withdrawn due to unhinged responses.

One advantage that OpenAI had over Google was having developed RLHF as a way to "align" the model's output to be more acceptable.

Part of Google's dropping the ball at that time period (but catching up now with Gemini) may also have been just not knowing what to do with it. It certainly wasn't apparent pre-ChatGPT that there'd be any huge public demand for something like this, or that people would find so many uses for it in API form, and especially so with LaMBDA's behavioral issues.

Re: Google's First Tensor Processing Unit: Architecture

#184
post #135

Earlier quoted context omitted.

Seems you have not worked with ML workloads, but base your comment on "internet wisdom", or worse, business analysts (I am sorry if that's inaccurate). On GPUs, ML "just works" (inference and training) and are always order of magnitude faster than whatever CPU you have. TPUs work very well for some model architectures (old ones that they were optimized and designed for) and on some novel others can be actually slower…

>On GPUs, ML "just works" If you had worked with ML, you'd know that this is not true. It's actually more like the opposite. It also has nothing to do with the chips themselves. Things don't magically work "because GPU", they work because manufacturers spend the time getting their drivers and ecosystems right. That's why for example noone is using AMD GPUs for ML, despite them offering more compute per dollar on pape…

Probably bartwr is using "GPUs" to mean NVIDIA GPUs. Seeing as nobody uses AMD GPUs for it, that simplification seems OK.

Re: Google's First Tensor Processing Unit: Architecture

#185
post #103
post #96

Earlier quoted context omitted.

> It's the only credible competition NVidia has This is wrong, both AMD and Intel (through Habana) have GPUs comparable to H100s in performance.

Yes, but they don't have the custom kernels that CUDA has. TPUs do have some!

Huh? The reason why they're competitive with Nvidia is because they have custom kernels for all the popular models.

Re: Google's First Tensor Processing Unit: Architecture

#186
post #135

Earlier quoted context omitted.

Seems you have not worked with ML workloads, but base your comment on "internet wisdom", or worse, business analysts (I am sorry if that's inaccurate). On GPUs, ML "just works" (inference and training) and are always order of magnitude faster than whatever CPU you have. TPUs work very well for some model architectures (old ones that they were optimized and designed for) and on some novel others can be actually slower…

>On GPUs, ML "just works" If you had worked with ML, you'd know that this is not true. It's actually more like the opposite. It also has nothing to do with the chips themselves. Things don't magically work "because GPU", they work because manufacturers spend the time getting their drivers and ecosystems right. That's why for example noone is using AMD GPUs for ML, despite them offering more compute per dollar on pape…

> That's why for example noone is using AMD GPUs for ML

You're right, they are behind, but to say that nobody is using it, is not truthful. AMD HPC clusters are being used [0] and [1] for AI/ML.

The larger issue is that AMD has only been building HPC clusters for the last period of time. Now, with the release of MI300x, we have Azure and Oracle coming online with them now. Disclosure, my business is also building a MI300x super computer as well, with the express goal of enabling more access to developers.

[0] https://defensescoop.com/2023/08/23/navys-new-25m-supercompu...

[1] https://arxiv.org/abs/2312.12705

Re: Google's First Tensor Processing Unit: Architecture

#187
post #109

Earlier quoted context omitted.

The way I see, NVidia only has a few advantages ordered from most important to least: 1. Reserved fab space. 2. Highly integrated software. 3. Hardware architecture that exists today. 4. Customer relationships. but all of these aspects are weak in one way or another: For #1, fab space is tight, and NVidia can strangle its consumer GPU market if it means selling more AI chips at a higher price. This advantage is gone…

Don’t underestimate CUDA as the moat. It’s been a decade of sheer dominance with multiple attempts to loosen its grip that haven’t been super fruitful. I’ll also add that their second moat is Mellanox. They have state of the art interconnect and networking that puts them ahead of the competition that are currently focusing just on the single unit.

This moat is going to get paralleled over the next few years. First off Mellanox is unobtanium with 52+ week lead times.

GigaIO has a PCIe fabric solution that is a fraction of the cost of Mellanox and available today. This enables up to 64 GPUs to appear on a single system.

We're also seeing the ultraethernet stuff come online as well, but that'll have to wait for PCIe6.

Re: Google's First Tensor Processing Unit: Architecture

#188
post #119

Earlier quoted context omitted.

There could be an opposite avenue: ad-free Google Premium subscription with AI chat as a crown jewel. An ultimate opportunity to diversify from ad revenue.

There's not enough money in it, as Google's scale. Especially because the people who'd pay for Premium tend to be the most prized people from an advertiser perspective. And most people won't pay, under any circumstances, but they will click on ads which make Google money.

The low operating margin of serving a GPT-4 scale model sounds like a compelling explanation for why Google stayed out of it.

But then why did Microsoft put its money behind it? Alphabet's revenue is around $300bn, and Microsoft's is around $210bn which is lower but it is the same order of magnitude.

Re: Google's First Tensor Processing Unit: Architecture

#189

Given what seems to be an enormous demand for fab space, when Microsoft or Google create a proprietary chip and need it produced how do they get to the front of the line? Are they simple enough that "older outdated less in demand" fabs can produce them? I know Apple and Nvidea has a lock on a lot of fab space?

They operate on outdated fab's (roughly state of the art - 1)

https://en.wikipedia.org/wiki/Tensor_Processing_Unit#Product...

They also do have a serious presence/spend on things like HBM, semianalysis has some good pieces on this.

Re: Google's First Tensor Processing Unit: Architecture

#190
I listened to a talk by Jim Keller from Tens torrent and their different approach to making AI cores - 5 Risc V cores one core for loading data, one for uploading data and the rest dedicated to performing matrix operations.

He did mention Google's TPU and the fact it was like programming a VLIW and they had about 500 people dedicated to their compiler.

Post reply on HN