Live data from Hacker News

Google supercharges machine learning tasks with TPU custom chip

cloudplatform.googleblog.com

271–280 of 283 posts

Re: Google supercharges machine learning tasks with TPU custom chip

#271

A podcast I listen to posted an interview with an expert last week saying that he perceived that much of the interest in custom hardware for machine learning tasks died when people realized how effective GPUs were at the (still-evolving-set-of) tasks. http://www.thetalkingmachines.com/blog/2016/5/5/sparse-codin... I wonder how general the gains from these ASIC's are and whether the performance/power efficiency wins w…

I listen to the Talking Machines as well. Great podcast. Another question would be are the gains worth the cost of an ML-specific ASIC. GPUs have the entire, massive gaming industry driving the cost down. I suppose that as adoption of gradient-descent-based neural networks increases, it may be worth the cost in a similar way that GPUs are worth the cost. Then again, I have never implemented SGD on a GPU so I'm not aw…

> massive gaming industry driving the cost down.

Per-unit manufacturing cost scales logarithmically. Even a single batch of custom silicon on yesterday's technology is only $30K. This is one of the reasons there is so much interest in RISC-V; hardware costs are not the barrier-to-entry that they used to be.

So yeah, the gaming market pushes the per-unit price of GPUs down, but even an additional 2x reduction in rackspace and power will pay for itself at the right scale.

Re: Google supercharges machine learning tasks with TPU custom chip

#272
post #196

It is interesting that they would make this into an ASIC, provided how notoriously high the development costs for ASICs are. Are those costs coming down? If so life will get very hard for the FPGA makers of the world soon. It would be interesting to see what the economics of this project are. I.e., what are the development costs and costs per chip. Of course it is very doubtful I will ever get to see the economics of…

It's mainly a question of volume. ASICs have a big economy of scale, so the cost-per-chip goes down considerably once you go over a certain number of chips. Plus, there's all the NRE costs of an ASIC design over an FPGA design. Google probably figured they could use enough chips to make the cost of manufacturing ASICs worthwhile. I don't think FPGAs are going to be beat out by ASICs for low volume applications anytim…

But even a custom run of silicon (on yesterday's technology) will only set you back $30K. That's one of the reasons there is so much interest in RISC-V.

Re: Google supercharges machine learning tasks with TPU custom chip

#273
post #72

Earlier quoted context omitted.

> By revealing that AlphaGo was based on this hardware Interesting, as the nature/science paper made no mention of this, it was exclusively trained on GPUs.

Note that the version of AlphaGo that beat Fan Hui and was presented in Nature is significantly different from the Version that played Lee Sedol. Unless you believe that they didn't work on it for half a year.

This explains a lot. Going into the match, Sedol thought he could beat DeepMind but thought it might only be a couple of years until the technology outpaced him. We knew Fan Hui and other Go professionals were helping the team, but a massive speedup is always nice too.

It's a bit underhanded, however. IMHO, the player should be able to study recent games before the match. But this is pretty typical, there were similar late-stage improvements with Chinook (checkers) and Deep Blue (chess).

Re: Google supercharges machine learning tasks with TPU custom chip

#274

Earlier quoted context omitted.

Just because you understand a machine doesn't mean it can't be dangerous. I could completely understand every aspect of a nuclear bomb, and I could still make a mistake and cause quite a bit of damage with a real one. Complex systems are notorious for having all sorts of unexpected problems, and mistakes happen all the time. How much complex software is entirely bug free? The danger of AI is more than just a random b…

You can condition the AI on the well being and freedom of human population. Hard to define precisely what that means, but it can be approximated with indirect measures. This is just what Asimov thought of in his novels. Another way to protect against catastrophe would be to launch multiple AI agents that optimize for the goal of nurturing humanity. They can keep each other in check. Also, humans will evolve as well.…

We don't know how to "condition" an AI to respect the well being and freedom of humanity. It's an extremely complicated problem with no simple solutions. Making an AI that wants to destroy humanity, however, is quite straightforward. Guess which one will most likely be built first?

Building multiple AIs doesn't solve anything. They can just as easily cooperate to destroy humanity as to help it.

Uploading humans won't be possible until we can already simulate intelligence in computers. We can't have uploads before AI.

Re: Google supercharges machine learning tasks with TPU custom chip

#275
post #260
post #229

I think the confluence of new technologies, and the re-emergence / rediscovery of older technologies is going to be the best combination. Whether it goes that way is not certain, since the best technology doesn't always win out. Here, though, the money should, since all would greatly reduce time and energy in mining and validating: * Vector processing computers - not von Neumann machines [1]. * Array languages new, o…

At this point in time, I think that the Python/Numpy stack offers the best performance, productivity, and expressiveness trade-off. With the [Numba]( http://numba.pydata.org ) just-in-time compiler, you can now easily bounce between numeric SIMD codes that leverage tuned BLAS/MKL, then go into more explicit loop-oriented constructs that perform equivalently to hand-coded C, while still being Python. If I were startin…

You may be right. I can't argue with Python's ubiquity; I have even steered my son in that direction, but with a hitch: I still had him learn some J.

The creator of Pandas, Wes McKinney, had a link up a few years back mentioning he was looking for people who were familiar with APL, J or K. It seems he was working on a new project/startup I think (could this have been the shuttered DataPad?). The links are dead now, but I will double check.

If the creator of Pandas is/was eyeing the older APL, and its newer brethren, I'd say it's a safe bet to keep J or K or Q on your radar because they fit. They're vector/array based; they are fast and iterative with a REPL; there is a lot of mathematical formalism in their origins and usage throughout the years, yet they are more beginner-friendly than say Haskell IMHO. I like Haskell too!

Re: Google supercharges machine learning tasks with TPU custom chip

#276
post #248

Earlier quoted context omitted.

I would suspect they're aiming for orders of magnitudes > 1k units though?

Unless they're doing 100K plus for 2-3 year life of a chip, it makes sense to stay w/ EasyPath or the Altera Hardcopy. It is much easier and more flexible for making revisions. Consider that they will most likely make another version, with new features in 2 years time. I'm sure the users of the chip will want that, just like any other system.

Altera Hardcopy has been discontinued.

https://www.altera.com/products/general/asic.highResolutionD...

> Altera no longer offers HardCopy structured ASIC products for new design starts. Altera continues to support HardCopy for existing designs.

Re: Google supercharges machine learning tasks with TPU custom chip

#277

Earlier quoted context omitted.

What do you think about http://optalysys.com/ or http://lighton.io ?

I think Optalysys looks interesting! For the curious, Optalysys has built a general purpose optics-based correlation/pattern matching machine. From some of their predecessor-company marketing material: The correlator performs pattern matching on large data sets such as high-resolution images, providing a measure of similarity and relative position between objects within the input scene. This allows large images [and…

This video was very cool. Are there any IC's that can perform analog computing for neural networks on the market now? I'm picturing something like an FPGA but with a bunch of op amps that you can connect into summers or amplifiers.

If not, how would one practically implement an analog computer for neural network programming (without several tables full of op-amps?)

Re: Google supercharges machine learning tasks with TPU custom chip

#278

Earlier quoted context omitted.

Because there's an absolute sippy straw of bandwidth to the thing if that's a 1x pci-e connection. For if it were delivering performance on par with a $1000 Maxwell class GPU, why wouldn't you guys crow about it? That would be a really big deal wouldn't it? TitanX for 20W? That'd be awesome. And having suffered through multiple pitches for us to buy various FPGA and boutique processors, I have yet to see someone who…

Hmm, I'd expect that once you've got your weights loaded machine learning would need much less bandwidth/flop than graphics does. Is that incorrect?

Apples and oranges. And as it turned out, it's an 8-bit fixed point processor. Useless for training without heroic handholding, limited use for inference. Google's PR machine does it again.

Re: Google supercharges machine learning tasks with TPU custom chip

#279
post #245
post #172

Earlier quoted context omitted.

"So"? If you're evaluating something today, how does it change your decision that we were late to market with Compute Engine (and in this specific case "bring-your-own-kernel")? If it's about future boring stuff, I think the list of boring stuff isn't too long ;). Disclosure: I work on Compute Engine.

"So" Google lost potential business for a while from people who wanted to spin up VMs rather than wanting to ship code to a proprietary execution framework.

I think you're agreeing: We certainly missed a huge segment of the market at the time, but now that we've got GCE new business can certainly come our way.

Re: Google supercharges machine learning tasks with TPU custom chip

#280
post #172

Earlier quoted context omitted.

"So"? If you're evaluating something today, how does it change your decision that we were late to market with Compute Engine (and in this specific case "bring-your-own-kernel")? If it's about future boring stuff, I think the list of boring stuff isn't too long ;). Disclosure: I work on Compute Engine.

All given, the fact that Google itself doesn't extensively use GC is kind of a red flag(I know quite a few Googlers from search infrastructure and none of them said their teams used GCE internally). A solid guarantee with AWS is if AWS goes down, then a multitude of Amazon's services also will go down(ex Amazonian myself), so it gives me a belief that AWS's uptime is more important to Amazon itself that it is for ext…

Search Infra (and Ads for that matter) is an extreme case. Google Search might be one of the worlds most highly tuned infrastructure projects: a marriage of code and hardware design to maximize performance, scoring, relevance and ultimately ROI.

Before we had custom machine types (November 2015 GA), we wouldn't have been remotely close to what they needed. I'm not even sure we've had anyone evaluate the amount of overhead KVM adds in either latency or throughput.

tl;dr: Don't let Search be your "not until they do it". We've got folks in Chrome, Android, VR, and more building on top of Cloud (as well as much of our internal tooling being on App Engine specifically).

Post reply on HN