Live data from Hacker News

Which GPU(s) to Get for Deep Learning

timdettmers.com

91–100 of 119 posts

Re: Which GPU(s) to Get for Deep Learning

#91
I'm wondering about two things:

1. Can laptops with, say NVidia 1070 or 1080 GPUs keep them cool at their stock frequencies for a few hours? My work laptop with non-U i7 (Thinkpad T440p) starts thermal throttling in just a few minutes when I compile something large.

2. Wouldn't two NVidia 1060 or 1070 outperform a single 1080 for training, assuming batch size is kept low enough so that each batch fits in a single card's memory?

Re: Which GPU(s) to Get for Deep Learning

#92
post #16

Earlier quoted context omitted.

Ok - p100's? The thing is that a K80 costs.... £4k what is it that people are paying for?

ECC memory (not that everyone really need that)...

In a tensorflow job I find it hard to imagine that a memory error would make a jot of difference... wouldn't the error simply be removed by the network training? If the error was in the execution then I imagine that the epoch would fail and it would then just be a matter of restarting.

Re: Which GPU(s) to Get for Deep Learning

#93
post #31

Earlier quoted context omitted.

If you're running a rackmount server, you need the Tesla series. We found that the GeForces tend to burn out when under heavy load, whereas we've not had a single Tesla series ever burn out.

Ooh ooh tell us more. Which GeForces and which server chassis? Adequate power supply?

Yes - more details please.

We've been using Titan-X and 1080 Maxwells in some Broadberry 4u chassis for the last year/18mths and we've had no burnouts so far.

I'm buying replacement pascals, and I can't justify teslas when I can get 8 geforces for the price of 1...

Re: Which GPU(s) to Get for Deep Learning

#95
post #47

Earlier quoted context omitted.

You might be interested in knowing that OpenCL is merging with Vulkan's API[1], so there'll likely be a lot more support in the future. [1]: https://news.ycombinator.com/item?id=14383296

The problem with OpenCL is the lack of well performing libraries, AMDs Core Math and Performance libraries are utter garbage still. If you can't do what Blender did and write everything from scratch including your primatives you'll be much slower than CUDA. But there is also a cost to it the Blender Cycles OpenCL code is nearly 5 times as big as their CUDA code and Cycles on OpenCL is still not at a feature parity wi…

I haven't heard much about AMD's ACML and ArrayFire but I'm not surprised.

On the other hand, I think you're underestimating the fact that OpenCL is an open spec and therefore has support from the FOSS world. CUDA has always been criticized for being closed source.

Even though OpenCL doesn't get much praise, and actually get criticized a bit (I'm personally not a big fan because writing OpenCL code is like bending over backwards). It has a lot of potential with new libraries coming up[1] which can possibly make it atleast on par with CUDA. Once that happens, AMD's cost effective cards will put it in the race.

[1]:https://github.com/clMathLibraries

Re: Which GPU(s) to Get for Deep Learning

#96
Short answer is that you can save weeks of work putting together a machine by just buying from Lambda (We of course use the 1080Tis):

https://lambdal.com/deep-learning-devbox

On average it takes a SWE or DL engineer a few days to set up a unit from scratch. Your company probably burns over $2,000/day so every day your DL engineer or SWE isn't up and running costs you money.

Re: Which GPU(s) to Get for Deep Learning

#97

Question: if I'm learning about neural networks and want to e.g. train a network to recognize MNIST digits, do I need a discrete graphics card (probably attached to a VPS that I would rent)? Or can I use the i5 Kaby Lake (which has an integrated GPU) in my laptop to train my network?

You can buy a 1 GPU machine with an i5 for $2,450 from Lambda Labs.

Re: Which GPU(s) to Get for Deep Learning

#98
post #93

Earlier quoted context omitted.

Ooh ooh tell us more. Which GeForces and which server chassis? Adequate power supply?

Yes - more details please. We've been using Titan-X and 1080 Maxwells in some Broadberry 4u chassis for the last year/18mths and we've had no burnouts so far. I'm buying replacement pascals, and I can't justify teslas when I can get 8 geforces for the price of 1...

Dell R720 / R730s, dual GPU (typically with K40m or K80) in there with 1100W dual redundant PSUs and the GPU enablement kit. We also set our fans to constant 70% min, to keep the airflow good.

On some servers we introduced the GTX 1080, either along-side a K40/K80 or two per chassis (see http://arnon.dk/how-does-the-nvidia-gtx-1080-stack-up-agains...).

They actually work 15% faster on average compared to the Tesla K series (Remember it's a 5 year old card), but they just stop working after a few months, or return inconsistent results for some operations.

Now, we're not doing graphics with them. We have a GPU database called SQream DB - and we depend on the results to be correct. In the end, they didn't make a lot of sense for us to deploy in a production environment, so back to the Tesla series we went.

Re: Which GPU(s) to Get for Deep Learning

#99
post #95

Earlier quoted context omitted.

The problem with OpenCL is the lack of well performing libraries, AMDs Core Math and Performance libraries are utter garbage still. If you can't do what Blender did and write everything from scratch including your primatives you'll be much slower than CUDA. But there is also a cost to it the Blender Cycles OpenCL code is nearly 5 times as big as their CUDA code and Cycles on OpenCL is still not at a feature parity wi…

I haven't heard much about AMD's ACML and ArrayFire but I'm not surprised. On the other hand, I think you're underestimating the fact that OpenCL is an open spec and therefore has support from the FOSS world. CUDA has always been criticized for being closed source. Even though OpenCL doesn't get much praise, and actually get criticized a bit (I'm personally not a big fan because writing OpenCL code is like bending ov…

[1] is actually AMD's old libraries. Judging by dev activity, it seems to be abandonware in favor of https://github.com/RadeonOpenCompute which is itself not ready for prime time.

Re: Which GPU(s) to Get for Deep Learning

#100
post #73
post #68

Earlier quoted context omitted.

The problem is that a month of GPU time on a cluster can buy you the hardware itself. If you are doing serious deep learning work it's just not cost effective at this point. If the costs come down by half or more, it may start looking viable for people who need a lot of resources.

Can you explain your math? A K80 on GCE is $.7/hr x 730 => $511/month if you were really 24x7. A K80 (and really we sell them by the die not the board) is more than $1000. I don't disagree that a consumer board is about that price, but they're not apples to apples. (Either in memory size, reliability or both). I'm fine with that being the real complaint: (major) cloud providers only sell the Tesla class boards, and t…

> but they're not apples to apples

It would be nice to see Nvidia or someone expand on this, so that users who have to make this choice can do so without guessing. If Google or AWS or M$ could publish reliability information, that'd be cool too.

Illustrative case: I run Monte Carlo work on GPUs and administer a local compute cluster. I tested a workload on a 16 GB P100 and a GTX 1080. A 12 GB P100 costs (academic, EU) 5000 euros while the latter costs 700 euros, but the performance difference is about 2x. Still, when we ask Nvidia reps, they say not to bother installing GTX cards in our cluster, because they aren't designed for 24/7 work, not commercializable etc. Even so, the GTX would have to burn out 3 times before the choice of P100 breaks even. Burning out three times means GTX 1080, then 1180, 1280 etc.

Post reply on HN