Live data from Hacker News

Comparing Google’s TPUv2 against Nvidia’s V100 on ResNet-50

blog.riseml.com

111–120 of 132 posts

Re: Comparing Google’s TPUv2 against Nvidia’s V100 on ResNet-50

#111

Hi, author here. The motivation for this article came out of the HN discussion on a previous post ( https://news.ycombinator.com/item?id=16447096 ). There was a lot of valuable feedback - thanks for that. Happy to answer questions!

I am not an ML guy, so I'm asking from a position of ignorance. (-:

But what's going on when some of the implementations of a standard algorithm don't converge, and different hardware has different accuracy rates on the same algorithm? Are DNNs really that flaky? And does it really make sense to be doing performance comparisons when the accuracy performance doesn't match?

Is the root problem that ResNet-50 works best with a smaller batch size?

And how do you do meaningful research into new DNNs if there's always an "Maybe if I ran it again over there I'd get better results" factor?

Thank you.

Re: Comparing Google’s TPUv2 against Nvidia’s V100 on ResNet-50

#112

Earlier quoted context omitted.

Not sure why I could not reply to your post so will reply here. Find the questioning on the entire stack just baffling with Google. People - Google has the strongest team of AI experts in the industry by a wide margin. At NIPS this year Google had more papers excepted than anyone else. The big AI breakthroughs come from Google. They solved Go a decade earlier than anyone thought possible. Would put FB #2 with AI expe…

Nobody is arguing Google is the best at AI. You are arguing that translates to Google is the best at making chips. They aren't, and there's no evidence they are. Tensorflow is open source, and Nvidia contributes. Would you be willing to place a bet that more people run tensorflow on tpu or GPU? Edit: and there are FAR more cuda users in general than tensorflow if you're trying to compare apples and oranges.

Well we have a data point that suggests they created a better chip. But just makes sense as they do the entire stack and that gives you the info they need to build a better chip.

Look a capsule networks and dynamic routing. That potentially drives a different architecture and Google has thousands of production models to use to optimize that Nvidia just does not have.

Plus it is one company so no IP issues.

But the biggy is we can see half the cost.

Re: Comparing Google’s TPUv2 against Nvidia’s V100 on ResNet-50

#113

Earlier quoted context omitted.

Nobody is arguing Google is the best at AI. You are arguing that translates to Google is the best at making chips. They aren't, and there's no evidence they are. Tensorflow is open source, and Nvidia contributes. Would you be willing to place a bet that more people run tensorflow on tpu or GPU? Edit: and there are FAR more cuda users in general than tensorflow if you're trying to compare apples and oranges.

Well we have a data point that suggests they created a better chip. But just makes sense as they do the entire stack and that gives you the info they need to build a better chip. Look a capsule networks and dynamic routing. That potentially drives a different architecture and Google has thousands of production models to use to optimize that Nvidia just does not have. Plus it is one company so no IP issues. But the bi…

No, you have a data point that says they created a chip that performed better on a single test for a single domain of work. You also have a data point that says tpu can't do any 64-bit simulations. It's not half the cost. See previous comment. It's about 100x the cost after 51 days.

Re: Comparing Google’s TPUv2 against Nvidia’s V100 on ResNet-50

#114

Earlier quoted context omitted.

If that 2x price/performance scales for all of Google's inferencing then it is definitely not a loss for them. If they can halve their running costs for inferencing then they are saving themselves a ton of money. Their TPUv2 was announced slightly before the V100 and the money savings they make by not paying Nvidia premiums probably helps. From the customer point of view, what is a GPU other than a specialised accele…

I agree, but chip development is an expensive business. There is nothing preventing Nvidia from immediately turning around and building a specialised ML accelerator with better software integration and higher bandwidth. For all we know they could already be working on one.

They already did two generations. Google has over $100B in the bank with less than $4B debt. So money is not an issue. It is tiny in the scheme of things.

Google has an advantage as they do the entire stack and can better optimize like we see here with half the cost.

Re: Comparing Google’s TPUv2 against Nvidia’s V100 on ResNet-50

#115

Earlier quoted context omitted.

Well we have a data point that suggests they created a better chip. But just makes sense as they do the entire stack and that gives you the info they need to build a better chip. Look a capsule networks and dynamic routing. That potentially drives a different architecture and Google has thousands of production models to use to optimize that Nvidia just does not have. Plus it is one company so no IP issues. But the bi…

No, you have a data point that says they created a chip that performed better on a single test for a single domain of work. You also have a data point that says tpu can't do any 64-bit simulations. It's not half the cost. See previous comment. It's about 100x the cost after 51 days.

Find the questioning on the entire stack just baffling with Google.

People - Google has the strongest team of AI experts in the industry by a wide margin. At NIPS this year Google had more papers excepted than anyone else. The big AI breakthroughs come from Google. They solved Go a decade earlier than anyone thought possible. Would put FB #2 with AI experts but a very distant #2.

Google miles ahead with SDC.

Plus Google is able to attract the top talent better than anyone else.

https://unsupervisedmethods.com/nips-accepted-papers-stats-2....

Applications - Search, Photos, Speech, AlphaZero, Self Driving Cars, Google now has over 4k NN in production. Nobody else even in the ball park. Hand down the leader in applications.

Infrastructure - Tensor Flow now has 98k stars on GitHub. It is the canonical AI framework in the industry and really nothing else close. CNTK is #2 with 14k stars. But Google ads about 7x per days stars.

https://github.com/tensorflow/tensorflow

Then there is Google cloud infrastructure and their other engineering talents which are well ahead of anyone.

I can go on but this is so incredibly silly. There is little question that Google is leading at every layer of the AI stack by a wide margin.

This is so silly I suspect something else going on here. We do not seem to be discussing things based on reality.

Is this about Damore?

BTW, Nvidia can only spin up what they know about. Google does not share everything but luckily they do a lot for Nvidia.

In 2018 you just have to run the infrastructure to be long term viable in the chip game.

Re: Comparing Google’s TPUv2 against Nvidia’s V100 on ResNet-50

#116

Earlier quoted context omitted.

Well we have a data point that suggests they created a better chip. But just makes sense as they do the entire stack and that gives you the info they need to build a better chip. Look a capsule networks and dynamic routing. That potentially drives a different architecture and Google has thousands of production models to use to optimize that Nvidia just does not have. Plus it is one company so no IP issues. But the bi…

No, you have a data point that says they created a chip that performed better on a single test for a single domain of work. You also have a data point that says tpu can't do any 64-bit simulations. It's not half the cost. See previous comment. It's about 100x the cost after 51 days.

Well do you have any test or seen any that indicates otherwise?

What we have here is same work and half the cost. Which makes sense as Google does all the layers of the stack and is just has a fundmental advantage to better optimize.

AI it is even more important.

Re: Comparing Google’s TPUv2 against Nvidia’s V100 on ResNet-50

#117

Earlier quoted context omitted.

Well we have a data point that suggests they created a better chip. But just makes sense as they do the entire stack and that gives you the info they need to build a better chip. Look a capsule networks and dynamic routing. That potentially drives a different architecture and Google has thousands of production models to use to optimize that Nvidia just does not have. Plus it is one company so no IP issues. But the bi…

No, you have a data point that says they created a chip that performed better on a single test for a single domain of work. You also have a data point that says tpu can't do any 64-bit simulations. It's not half the cost. See previous comment. It's about 100x the cost after 51 days.

Also have no idea what 100x the cost refers to. We can see half the cost.

Re: Comparing Google’s TPUv2 against Nvidia’s V100 on ResNet-50

#118

Earlier quoted context omitted.

I agree, but chip development is an expensive business. There is nothing preventing Nvidia from immediately turning around and building a specialised ML accelerator with better software integration and higher bandwidth. For all we know they could already be working on one.

They already did two generations. Google has over $100B in the bank with less than $4B debt. So money is not an issue. It is tiny in the scheme of things. Google has an advantage as they do the entire stack and can better optimize like we see here with half the cost.

Nvidia is actively building an entire deep learning stack internally, all the way to releasing a self-driving simulation platform which they are using to build their own self-driving software [1].

I think they are actually farther along and more aggressive about exploring deep learning use cases in production than Google today; augmenting real data with extensive simulation is really a far-reaching idea that comes directly from their gaming experience.

> So money is not an issue. It is tiny in the scheme of things.

Money of course is always an issue long term; otherwise why doesn't Google Fiber just spend tens of billions of dollars to build out its nationwide network? Because it will see negative ROI even if they succeed.

The TPU has to eventually make a real return to Google, and it won't if nvidia can spend the same amount of money and build a faster product and sell it to all the other cloud players, which I believe they definitely can.

Put another way, the TPU has to be cheaper to Google than buying nvidia GPUs after factoring in its development costs, whereas nvidia gets to amortize those dev costs over all other cloud providers and all other GPU customers. Google isn't about to sell the TPU to other cloud providers; the entire idea is to use it to drive Google Cloud adoption.

The TPU is a fine chip, but if you just look at the big picture, there is every sign that nvidia could build the same or better product for less money because it has far more synergies across the hardware and chip design stack; e.g. the TPU only has PCIe connectors, while nvidia has already worked with IBM to get NVLink into supercomputers [2]. For some workloads the TPU will likely be bandwidth-starved communicating with the CPU and main memory.

[1] https://nvidianews.nvidia.com/news/nvidia-introduces-drive-c...

[2] https://www.ibm.com/us-en/marketplace/power-systems-ac922/de...

Re: Comparing Google’s TPUv2 against Nvidia’s V100 on ResNet-50

#119

Earlier quoted context omitted.

They already did two generations. Google has over $100B in the bank with less than $4B debt. So money is not an issue. It is tiny in the scheme of things. Google has an advantage as they do the entire stack and can better optimize like we see here with half the cost.

Nvidia is actively building an entire deep learning stack internally, all the way to releasing a self-driving simulation platform which they are using to build their own self-driving software [1]. I think they are actually farther along and more aggressive about exploring deep learning use cases in production than Google today; augmenting real data with extensive simulation is really a far-reaching idea that comes di…

The problem is Nvidia is never going to have the AI expertise up and down the stack like Google.

As far as I am aware Nvidia does not even run a cloud do they? Obviously never going to have the production NN that Google has.

Google now has well over 4k NN in production and not sure if Nvidia has any? Well over a billion a day are using the Google NN. That data allows Google to iterate in ways that Nvidia just never would be able to.

But this was all theory and why starting to see a little more concrete results like this where Google with their TPUs able to charge 1/2 the price of using Nvidia is value. Then we also have the paper from Google on the Gen 1.

I would guess Google is working on a gen 3. Nvidia is trying to catch a moving target but without the data. So they are behind, trying to catch up, but missing an arm.

A perfect example of this phenomenon is Capsule network pioneered by Hinton. They use dynamic routing which is potentially going to require different approach to memory access as the pattern would be different than CNN or RNN.

Today the problem is memory access and no longer instruction execution. Google nailed the low hanging fruit with the Gen 1 TPUs. They have 65536 very simple cores. Now you have to go after memory access.

Your post is all over the place so a bit hard to respond. Google Fiber was NOT about cost. It was about AT&T and other established players with some local governments making it difficult for Google to access what they needed to be able to compete.

I hate debating something with someone that is doing what you are doing. Google Fiber? Really?

"I think they are actually farther along and more aggressive about exploring deep learning"

I do a LOT of surfing on sites and can easily say this is the craziest thing I have read in a bit. You are honestly comparing Nvidia to Google? Really?

Google solved Go a decade early. Hinton did the Capsule networks and basically the farther of DL. Well made it actually work. What breakthrough came from Nvidia?

A single one?

There is so much crazy stuff in your posts this must be driven by something else and something emotional? Your points are just not based on reality. Is this really about Google firing Damore?

BTW, Nvidia read the Google Gen 1 TPU paper and why we see them doing similar things. But Google is going to move to addressing the memory access problems as that is the next area to improve. Once Google figures it out then you will see Nvidia just copy the approach like they are doing with the gen 1 TPUs.

I listened to this Nvidia presentation on YouTube and they were basically quoting the Google TPU paper. Talking about using 8 bit, integers, etc, for inference.

Google will release the gen 3 and then share a paper on the gen 2 and we will see Nvivida then try to copy that one. Nvidia always a couple of steps behind.

But I am a super curious person and can you share what this is really all about?

Re: Comparing Google’s TPUv2 against Nvidia’s V100 on ResNet-50

#120

Earlier quoted context omitted.

No, you have a data point that says they created a chip that performed better on a single test for a single domain of work. You also have a data point that says tpu can't do any 64-bit simulations. It's not half the cost. See previous comment. It's about 100x the cost after 51 days.

Also have no idea what 100x the cost refers to. We can see half the cost.

This isn't worth going on about anymore. If you can't the cost of running a GPU in your own server versus the TPU in the cloud, then I'm not going to help you.
Post reply on HN