The Nvidia DGX-1 Deep Learning Supercomputer in a Box
91–100 of 106 posts
Re: The Nvidia DGX-1 Deep Learning Supercomputer in a Box
#92I am looking forward to OpenCL catching up with CUDA in maturity and adoption, so that NVidia's monopoly in Silicon for deep learning will come to an end.
Having C only wasn't a good idea. NVidia was quite clever in giving first class treatment to C++, Fortran and any compiler vendor that wished to target PTX.
Also the visual debugging tools are quite good.
Khronos apparently needed to be hit hard to realise that not everyone wants to be stuck with C for HPC in the 21st century.
Also although Apple is the creator of OpenCL, they don't seem to give much love to it.
Then you have Google caring about it's Renderscript dialect, which doesn't help to the overall uptake in OpenCL.
There isn't a monopoly, rather vendors that lacked the perception to appeal to the developers wanted to have as tooling and performance.
Anyone is free to go use OpenCL, use C or a language with a compiler with a C target, do printf debugging and feel free.
Are any vendors already doing SPIR support?
Re: The Nvidia DGX-1 Deep Learning Supercomputer in a Box
#93Looks like a research in machine learning will only be done in huge corporations. You'll need an amount of funding comparable to LHC. Time to use better models like kernel ensembles, maybe they are not that accurate, but they are easier to train on a single CPU.
Re: The Nvidia DGX-1 Deep Learning Supercomputer in a Box
#94Just for some perspective, a little over 10 years ago, this $130k turnkey installation would sit at #1 in TOP500, easily beating out hundred-million-dollar initiatives like NEC's Earth Simulator and IBM's BlueGene/L: http://www.top500.org/lists/2005/06/ (170 TFLOPS vs. 137 TFLOPS) At the other end, even a single GTX 960 would make it onto the list, placing in the 200s.
The 170 TFLOPs number that NVIDIA gives out is for FP16, while the Top 10 list gives its number for for FP64. The P100 that makes up this NVIDIA box gives about 5.3TFLOPs per card, or a total of 42.4TFLOPs for the whole box. Sure, you can say that deep learning doesn't need FP64, but it is REALLY unfair to compare this to anything on the TOP500 list, especially when you consider the fact that this is not balanced in…
Re: The Nvidia DGX-1 Deep Learning Supercomputer in a Box
#95Re: The Nvidia DGX-1 Deep Learning Supercomputer in a Box
#96Earlier quoted context omitted.
If that sounds insane, you're going to lose your mind when you realize how many KILOwatts your oven uses. 3.2KW is less than a dishwasher.
Are you sure your numbers are right? What kind of dishwasher do you have? And what kind of oven? For the US, at least, most dishwashers are well under 1600W, and few ovens exceed under 3200W. https://www.daftlogic.com/information-appliance-power-consum...
Re: The Nvidia DGX-1 Deep Learning Supercomputer in a Box
#97Just for some perspective, a little over 10 years ago, this $130k turnkey installation would sit at #1 in TOP500, easily beating out hundred-million-dollar initiatives like NEC's Earth Simulator and IBM's BlueGene/L: http://www.top500.org/lists/2005/06/ (170 TFLOPS vs. 137 TFLOPS) At the other end, even a single GTX 960 would make it onto the list, placing in the 200s.
The 170 TFLOPs number that NVIDIA gives out is for FP16, while the Top 10 list gives its number for for FP64. The P100 that makes up this NVIDIA box gives about 5.3TFLOPs per card, or a total of 42.4TFLOPs for the whole box. Sure, you can say that deep learning doesn't need FP64, but it is REALLY unfair to compare this to anything on the TOP500 list, especially when you consider the fact that this is not balanced in…
*http://www.anandtech.com/show/10222/nvidia-announces-tesla-p...
Re: The Nvidia DGX-1 Deep Learning Supercomputer in a Box
#98Anyone have any idea of how the GPUs in this machine compare to the GPUs in their high end gaming products?
Re: The Nvidia DGX-1 Deep Learning Supercomputer in a Box
#99Earlier quoted context omitted.
The 170 TFLOPs number that NVIDIA gives out is for FP16, while the Top 10 list gives its number for for FP64. The P100 that makes up this NVIDIA box gives about 5.3TFLOPs per card, or a total of 42.4TFLOPs for the whole box. Sure, you can say that deep learning doesn't need FP64, but it is REALLY unfair to compare this to anything on the TOP500 list, especially when you consider the fact that this is not balanced in…
Still, how many of these boxes are we talking about to match the performance of the #1 from the top 500 in 2005? 10? 20? That's still under $3m for 20, which is pretty impressive to me.
Re: The Nvidia DGX-1 Deep Learning Supercomputer in a Box
#100$129k for this machine. In the keynote its interesting that they mentioned the product line being: "Tesla M40 for hyperscale, K80 for multi-app HPC, P100 for scales very high, and DGX-1 for the early adopters". The GP100/P100 with the 16nm process probably gives a considerable performance/power advantage over the Tesla... but this gives me the feeling that we may not see consumer or workstation-level Pascal boards fo…