Live data from Hacker News

The Nvidia DGX-1 Deep Learning Supercomputer in a Box

nvidia.com

101–106 of 106 posts

Re: The Nvidia DGX-1 Deep Learning Supercomputer in a Box

#101
post #15
post #9

Check out the specs here: http://images.nvidia.com/content/technologies/deep-learning/... though I'm most curious about what motherboard is in there to support NVLink and NVHS. Good overview of Pascal here: https://devblogs.nvidia.com/parallelforall/inside-pascal/ 1 question: will we see NVLink become an open standard for use in/with other coprocessors? 1 gripe: they give relative performance data as compared to a CP…

You mean you're not surprised that a machine with 8 GPUs, apparently costing $129k USD (from comment below), can outperform a single CPU? :) (Of course, a better metric is that it's getting ~56x the performance at probably ~10x the TDP, but that's not surprising for a GPU with the current state of deep learning code.) To their credit, the thermal and power engineering needed to get that dense a compute deployment is…

[deleted]

Re: The Nvidia DGX-1 Deep Learning Supercomputer in a Box

#102
I have to wonder about intel and their Xeon Phi range. Last I checked they were supposed to launch a followup late last year that never manifested. Now we're 4 months in 2016 and still no new phi's.

Couple that with the fact that they want you to use their compilers (extremely expensive), on a specialized system that can support the card, and you get a platform that nobody other than supercomputer companies can reasonably use. Meanwhile any developer who want to try something with cuda can drop $200 dollars on a GPU and go, then scale accordingly. I think intel somewhat acknowledged this by having a firesale on phi cards and dev licenses last year but it was only for a passively cooled model (really only works well in servers, not workstations).

Intel do this:

  - Offer a $200-400 XEON PHI CARD
  - Include whatever compiler needed to use it with the card
  - Make this easily buyable
  - Contribute ports of Cuda-based frameworks over to Xeon Phi
I feel like they could do this pretty easily, even if it lost money, it's pennies compared to what they're going to lose if nvidia keeps trumping them on machine learning. They need to give dev's the tooling and financial incentive to write something for Phi instead of cuda, right now it completely doesn't exist and frameworks basically use Cuda by default.

If you're AMD, do the same thing but replace the phrase Xeon Phi with Radeon/Firepro

Re: The Nvidia DGX-1 Deep Learning Supercomputer in a Box

#103
post #97

Earlier quoted context omitted.

The 170 TFLOPs number that NVIDIA gives out is for FP16, while the Top 10 list gives its number for for FP64. The P100 that makes up this NVIDIA box gives about 5.3TFLOPs per card, or a total of 42.4TFLOPs for the whole box. Sure, you can say that deep learning doesn't need FP64, but it is REALLY unfair to compare this to anything on the TOP500 list, especially when you consider the fact that this is not balanced in…

Was thinking the same, but look at the memory bandwidth of this thing, 720GB/sec.* That is the number you should look at, and that is a sweet number... Also the NVLink tech looks nice for multi-GPU/heterogeneous computing (I guess the latency is the most important, but no idea how that is ). Do's some one know if the Xeons are connected to the GPUs with NVLink or are they on PCIE ? (I know the new POWER chips have NV…

PCIe is the huge bottleneck... The two Xeon's in the box are 2698v3's, which each have only 40 PCIe lanes, meaning they are restricted to using 8x PCIe3 lanes per card, which would net you a whopping 8GB/s between each CPU and GPU. EDIT: Oh, and no, Intel does and probably never will support NVLink. I will eat my words (type this post up on paper and eat it) if they do in the next 5 years.

When I talk about balanced (which is a huge influence in my architectural and system level designs), I want to ideally be able to hit theoretical throughput. If we look at FP64 as an example, if I want to have sustained throughput of fused multiply adds (which is how NVIDIA always advertises their theoretical FLOP numbers as), I would be needing to move 196 data bits (three 64 bit floating point operands) in to each of my FPUs every cycle, and 64 bits out. 256 bits per cycle in a fully pipelined situation to be able to do 2 FLOPs/cycle. So if our ideal bandwidth is 16 Bytes for every 1 FLOP, if you have almost 10x more floating point capability than memory bandwidth, you are going to have a bad time (and GPUs very well reflect this on memory intensive workloads... take a look at GPUs on HPCG, they only get ~1-3% of their theoretical peak).

I'm working on my own HPC targeted chip, so obviously have some bias there, but 720GB/s memory bandwidth for a chip that is that large and using that much power isn't that impressive to me. Obviously I should wait to boast until I have my silicon in hand, but getting more than 3/4ths of that bandwidth in less than 1/10th of the power. Add in some fancy tricks and our goal is having our advertised theoretical numbers be pretty damn close to real application performance for memory intensive workloads.

Re: The Nvidia DGX-1 Deep Learning Supercomputer in a Box

#104
post #97

Earlier quoted context omitted.

Was thinking the same, but look at the memory bandwidth of this thing, 720GB/sec.* That is the number you should look at, and that is a sweet number... Also the NVLink tech looks nice for multi-GPU/heterogeneous computing (I guess the latency is the most important, but no idea how that is ). Do's some one know if the Xeons are connected to the GPUs with NVLink or are they on PCIE ? (I know the new POWER chips have NV…

PCIe is the huge bottleneck... The two Xeon's in the box are 2698v3's, which each have only 40 PCIe lanes, meaning they are restricted to using 8x PCIe3 lanes per card, which would net you a whopping 8GB/s between each CPU and GPU. EDIT: Oh, and no, Intel does and probably never will support NVLink. I will eat my words (type this post up on paper and eat it) if they do in the next 5 years. When I talk about balanced…

>PCIe is the huge bottleneck... The two Xeon's in the box..

It's a wast to put Xeon's on this things if they use the PCIe, you end up in a loot of cases only using them to drive the GPU's.

>When I talk about balanced (which is a huge influence in my architectural and system level designs)...

The DP performance on Tesla's is ridiculous, think it is a marketing ploy. People talk of buying gaming cards.., as you are almost always memory bound..

>I'm working on my own HPC targeted chip..

Looks nice, you are throwing out all HW bloat and doing everything in software? Are you planing to have some form of OS running on this chips?

Re: The Nvidia DGX-1 Deep Learning Supercomputer in a Box

#105
post #74

Earlier quoted context omitted.

Are you sure your numbers are right? What kind of dishwasher do you have? And what kind of oven? For the US, at least, most dishwashers are well under 1600W, and few ovens exceed under 3200W. https://www.daftlogic.com/information-appliance-power-consum...

Yeah, 3.2 KW would mean the dishwasher's heating the dishes to high temperatures, so it'd be more of an oven than a dishwasher.

Dishwashers not only heat the water, they also heat up the entire compartment if you have the drying feature turned on. They are basically an oven.

You can cook in them as well: http://www.thekitchn.com/can-you-really-cook-salmon-in-a-dis...

Re: The Nvidia DGX-1 Deep Learning Supercomputer in a Box

#106

Just for some perspective, a little over 10 years ago, this $130k turnkey installation would sit at #1 in TOP500, easily beating out hundred-million-dollar initiatives like NEC's Earth Simulator and IBM's BlueGene/L: http://www.top500.org/lists/2005/06/ (170 TFLOPS vs. 137 TFLOPS) At the other end, even a single GTX 960 would make it onto the list, placing in the 200s.

already mentioned 16FP 170 TFLOPS (that is 64 FP 42.5TFLOPS) of DGX-1. There is also issue of GPU vs CPU: basically you couldn't directly compare these operations on same scale. You could easily drop 100x of your GPU performance at bad case scenario. Basic idea of GPU that you could possible gain sometimes extra 1000s times performance
Post reply on HN