Live data from Hacker News

The Nvidia DGX-1 Deep Learning Supercomputer in a Box

nvidia.com

71–80 of 106 posts

Re: The Nvidia DGX-1 Deep Learning Supercomputer in a Box

#71
post #15

Earlier quoted context omitted.

You mean you're not surprised that a machine with 8 GPUs, apparently costing $129k USD (from comment below), can outperform a single CPU? :) (Of course, a better metric is that it's getting ~56x the performance at probably ~10x the TDP, but that's not surprising for a GPU with the current state of deep learning code.) To their credit, the thermal and power engineering needed to get that dense a compute deployment is…

$129K buys you a lot of dual 22-core servers.

if you include space, power, HVAC, and networking?

Re: The Nvidia DGX-1 Deep Learning Supercomputer in a Box

#72
post #49

Earlier quoted context omitted.

Uh, using what hardware? The 980 Ti is about 11 TFLOP in half-precision (apples to apples). So 16x 980 Ti cards would take up twice as much rack space for $11k. Your estimate (and NVIDIA's pricing) is off by more than an order of magnitude...

A 980 Ti doesn't have FP16 hardware. The only Maxwell based component with such support is their Tegra part.

[deleted]

Re: The Nvidia DGX-1 Deep Learning Supercomputer in a Box

#73
post #61
post #49

Earlier quoted context omitted.

Uh, using what hardware? The 980 Ti is about 11 TFLOP in half-precision (apples to apples). So 16x 980 Ti cards would take up twice as much rack space for $11k. Your estimate (and NVIDIA's pricing) is off by more than an order of magnitude...

Isn't the ECC tax around 10x?

So, just do the computation twice and compare the results?

OK, so there's twice the power to pay for but it seems like at $129k acquisition cost per 3.2KW consumption you could run for tens of years before break-even.

Re: The Nvidia DGX-1 Deep Learning Supercomputer in a Box

#74
post #39

Earlier quoted context omitted.

To me, $129k isn't surprising since it is only going to be bought by researchers with big budgets. Small-timers will still build 3x GTX980 systems for under $5k. 3.2 KILOwatts sounded insane to me, but I suppose you'll have your own server rack to put it in if you can afford to buy one of these.

If that sounds insane, you're going to lose your mind when you realize how many KILOwatts your oven uses. 3.2KW is less than a dishwasher.

Are you sure your numbers are right? What kind of dishwasher do you have? And what kind of oven? For the US, at least, most dishwashers are well under 1600W, and few ovens exceed under 3200W.

https://www.daftlogic.com/information-appliance-power-consum...

Re: The Nvidia DGX-1 Deep Learning Supercomputer in a Box

#75
post #39

Earlier quoted context omitted.

To me, $129k isn't surprising since it is only going to be bought by researchers with big budgets. Small-timers will still build 3x GTX980 systems for under $5k. 3.2 KILOwatts sounded insane to me, but I suppose you'll have your own server rack to put it in if you can afford to buy one of these.

If that sounds insane, you're going to lose your mind when you realize how many KILOwatts your oven uses. 3.2KW is less than a dishwasher.

Forget that, the HVAC unit for my home is rated at 24.6kW.

Re: The Nvidia DGX-1 Deep Learning Supercomputer in a Box

#76
post #15

Earlier quoted context omitted.

You mean you're not surprised that a machine with 8 GPUs, apparently costing $129k USD (from comment below), can outperform a single CPU? :) (Of course, a better metric is that it's getting ~56x the performance at probably ~10x the TDP, but that's not surprising for a GPU with the current state of deep learning code.) To their credit, the thermal and power engineering needed to get that dense a compute deployment is…

$129K buys you a lot of dual 22-core servers.

What kind of server pricing are you getting? Base servers are cheap, but add high-end Xeons and memory, not to ignore interconnect and I get something like 7 ok configured 1U servers for $129K (2 20 core w lots of RAM, 10GbE NICs and mirrored boot/swap). No interconnect switching. That's for 20 core Haswell because I don't yet have discount pricing for Broadwell Xeons. I'm sure one could do better at hyperscaler discount but this is startup low-ish quantity.

Re: The Nvidia DGX-1 Deep Learning Supercomputer in a Box

#78

Just for some perspective, a little over 10 years ago, this $130k turnkey installation would sit at #1 in TOP500, easily beating out hundred-million-dollar initiatives like NEC's Earth Simulator and IBM's BlueGene/L: http://www.top500.org/lists/2005/06/ (170 TFLOPS vs. 137 TFLOPS) At the other end, even a single GTX 960 would make it onto the list, placing in the 200s.

That is 170 TFLOPS Rpeak (theoretical performance assuming you could find a workload that doesn't need to wait for data movement) at half precision vs. 137 TFLOPS Rmax (usable performance on a dummy linear algebra problem) at double precision. No, it would not top the list.

Re: The Nvidia DGX-1 Deep Learning Supercomputer in a Box

#79

Earlier quoted context omitted.

It looks like it uses a separate daughterboard that houses the GPUs + NVLink, connected to the main motherboard using quad Infiniband EDR (400Gbps) + RDMA. http://images.anandtech.com/doci/10225/SSP_85.JPG

The diagram is confusing, but the GPUs are connected to the NVLink matrix which is connected to the motherboard via the PLX PCIe switches. The quad IB/dual 10GbE are separate IO attached to the motherboard. https://devblogs.nvidia.com/parallelforall/inside-pascal/

That would make much more sense. Thanks! The PCI bandwidth must be fairly limited. 4x 100G Infiniband is 64x PCIe lanes, out of 80x lanes available.

Re: The Nvidia DGX-1 Deep Learning Supercomputer in a Box

#80
post #15

Earlier quoted context omitted.

You mean you're not surprised that a machine with 8 GPUs, apparently costing $129k USD (from comment below), can outperform a single CPU? :) (Of course, a better metric is that it's getting ~56x the performance at probably ~10x the TDP, but that's not surprising for a GPU with the current state of deep learning code.) To their credit, the thermal and power engineering needed to get that dense a compute deployment is…

$129K buys you a lot of dual 22-core servers.

Not really. Not a lot at least.
Post reply on HN