Live data from Hacker News

Radeon Instinct – Optimized Machine and Deep Learning

radeon.com

71–80 of 88 posts

Re: Radeon Instinct – Optimized Machine and Deep Learning

#71
post #68

Ahh, what exciting times we live in. Just look at the example applications: - autonomous vehicles - autopilot drone - personal assistant - personal robots - ... i know it's optimistic, but it's not science-fiction.

- running out of natural resources - child slavery to build new iPhones in Africa and China - killing and burning wildlife to build new farms in Latin America - still not having any solution for problem of drinkable water in 2/3 of World What a time to be alive!

well, i was very active for greenpeace/other environmental organisations/the green party youth in germany. I know of these issues and it's getting better, but it's still awfully slow and frustrating. If you want some improvements way fewer are starving[0], food security improved for millions. I know there is no rational reason for children to starve, humanity produces enough food, but frankly we (as in first world) don't care. We really don't care. Most just donate something on the end of the month and think they are no part of the solution. But try explaining to people that they should scale back their consumption. They do stuff as long as there is not the slightest inconvenience.

I really think that in a weird way technology can help (not solve) overcome, or better compensate for our human stupidity in some of these issues. Not alone, of course. But right now i don't see a path for me that's more effective in tackling these problems at the same rate as improving technology. (if you want to see it more pragmatic: NGO's really need some better tools, especially for campaigning and training (this may be the most important, most are volunteers). It is often very unprofessionlized and inefficient. It's really holding some of them back.) Maybe there will be a time when i see more potential impact for my personal political ambitions, but right now i don't see how this is going to work better than studying computer science.

Also, it's been worse for humanity ;) We dropped like flies in the end of the 19. century due to diseases.

[0] https://www.wfp.org/news/news-release/world-hunger-falls-und...

Re: Radeon Instinct – Optimized Machine and Deep Learning

#72
post #23

Earlier quoted context omitted.

The tricky, but fun parts. It's certainly not trivial to port those, but I'm not too worried about it as long as there is solid runtime support in ROCm. More fun work for perf engineers like me. ;) Massive amount of boilerplate, heavy APIs or crappy software stack are more dangerous IMO than having to drop in replacements for NVIDIA-specific optimizations of GPU-to-NIC or GPU-to-GPU communication.

Not only not trivial to port those but they are constantly will have to play catch up and be at the mercy of NVIDIA in regards to spec. AMD is in a catch 22, support CUDA and be effectively in a constant catch up position, not support it and continue to be ignored by the market at large simply because the momentum NVIDIA has managed to achieve with CUDA over the years.

There are several well known (certainly to AMD) strategies and tactics for dealing with this situation.

One example: AMD can support CUDA while focusing on price/performance at the low end (initiating an Innovators Dilemma for NVIDIA), and if successful then solidify AMD's position, for example by starting a standards process for a successor to CUDA.

NVIDIA, of course, has various countering moves available such as IP moats, initiating competing standards, and so on.

Rarely is a credible player entirely boxed in. And even then, there are moves available to make the best of things like making some acquisitions and then spinning off the division into an independent company, or selling it off to another player (in this case probably Intel, but maybe ARM could be suckered into it) and so on.

Re: Radeon Instinct – Optimized Machine and Deep Learning

#74
post #12

What's particularly interesting here is that the Fiji card they propose is a very different beast than any of the NVIDIA offerings. The MI8 card's HBM has a great power and performance advantage (512 GB/s peak bandwidth) even if it's on 28 nm. NVIDIA has nothing that has even remotely comparable bandwidth in this price/perf/TDP regime. None of the NVIDIA GP10[24] Teslas have GDDR5X -- not to surprising given that it…

I wonder if something like AMD's SSD+GPU combo would improve the performance of this? Increasing the usable buffer size.

AFAIK AMD does not produce SSDs.

Having said that there are some early techniques out there for DMA from NVMe drives to GPU ram directly. Like this: http://kaigai.hatenablog.com/entry/2016/09/08/003556

Re: Radeon Instinct – Optimized Machine and Deep Learning

#75
post #74

Earlier quoted context omitted.

I wonder if something like AMD's SSD+GPU combo would improve the performance of this? Increasing the usable buffer size.

AFAIK AMD does not produce SSDs. Having said that there are some early techniques out there for DMA from NVMe drives to GPU ram directly. Like this: http://kaigai.hatenablog.com/entry/2016/09/08/003556

The parent is referring to this. https://www.extremetech.com/extreme/232416-amd-announces-new...

Re: Radeon Instinct – Optimized Machine and Deep Learning

#76
post #3

Earlier quoted context omitted.

If AMD can release solid hardware, all they would need to do is add support for their hardware to popular open-source projects like TensorFlow. I just hope they do the second part correctly.

It looks like they're doing it, or already have done so. https://github.com/GPUOpen-ProfessionalCompute-Tools/HIP/iss...

TensorFlow is evolving rapidly and it's a challenge to keep up with GPU driver updates, CUDA updates, TensorFlow updates, and not fall into version hell.

For instance I took the Udemy course a few months back, given by a Google scientist, and in at least one instance the code was pinned to CPU because it didn't run on GPU, presumably due to a TensorFlow issue.

If they can get TensorFlow running well on AMD and release images that make it painless, it will be a huge win, but I'm a little skeptical. Big difference between getting a benchmark win and a stable platform.

Re: Radeon Instinct – Optimized Machine and Deep Learning

#77

Earlier quoted context omitted.

Not only not trivial to port those but they are constantly will have to play catch up and be at the mercy of NVIDIA in regards to spec. AMD is in a catch 22, support CUDA and be effectively in a constant catch up position, not support it and continue to be ignored by the market at large simply because the momentum NVIDIA has managed to achieve with CUDA over the years.

There are several well known (certainly to AMD) strategies and tactics for dealing with this situation. One example: AMD can support CUDA while focusing on price/performance at the low end (initiating an Innovators Dilemma for NVIDIA), and if successful then solidify AMD's position, for example by starting a standards process for a successor to CUDA. NVIDIA, of course, has various countering moves available such as I…

> if successful then solidify AMD's position, for example by starting a standards process for a successor to CUDA.

I thought OpenCL was the standard in this space.

Re: Radeon Instinct – Optimized Machine and Deep Learning

#78
post #74

Earlier quoted context omitted.

AFAIK AMD does not produce SSDs. Having said that there are some early techniques out there for DMA from NVMe drives to GPU ram directly. Like this: http://kaigai.hatenablog.com/entry/2016/09/08/003556

The parent is referring to this. https://www.extremetech.com/extreme/232416-amd-announces-new...

I was not aware of this. Thanks for correcting me.

Re: Radeon Instinct – Optimized Machine and Deep Learning

#79
post #38

Earlier quoted context omitted.

NVIDIA is playing catch-up with themselves and the shifting market too! Just look at the Maxwell-based Tesla cards that came out of the blue (as if they were an afterthought); I bet Facebook, Baidu, Google, etc. told NVIDIA that Kepler was shit for their use-cases and they did not want to wait for Pascal. Or look at the bizarre Pascal product line where the GP100 does support half precision, but the others don't, not…

Not sure if Maxwell came out of the blue, Maxwell 1 was designed for mobile, embedded and tesla, Maxwell 2 came out most likely because Pascal at large was delayed. As for the half precision, it's pretty much the same thing NVIDIA been doing since Kepler dumping FP64 and FP16, especially FP16 due to the silicon costs. NVIDIA came out with the Titan and Titan Black with baller FP16 performance and no one seem to care,…

> Not sure if Maxwell came out of the blue, Maxwell 1 was designed for mobile, embedded and tesla, Maxwell 2 came out most likely because Pascal at large was delayed.

First off not sure what you are referring to by "Maxwell 1" and "Maxwell 2"; there's GM204, GM206, GM200 all very similar, and GM20B the slight outlier.

Maxwell was an arch tweak on Kepler which, due to everyone but Intel stuck at 28 nm, had to cut down on all but the gaming-essential stuff to deliver what the consumers expected (>1.5x gen-boost). For that reason, and because HPC has traditionally been thought to need DP, IIRC no public roadmap mentioned Maxwell Tesla at all. Instead, the GK210 (K80) was the attempt to be the bridge-gap chip for HPC until Pascal; it was released just a few months before the big GM200 consumer chips came out. That is until fall '15 when the late arriving M40 and M4 were pitched as "Deep learning accelerators" (though GRID/virtualization version were released a little earlier in August). Quite obvious naming, plus no sane HPC shop would buy a crippled (no DP) chip > As for the half precision, it's pretty much the same thing NVIDIA been doing since Kepler dumping FP64 and FP16, especially FP16 due to the silicon costs.

Not really. In Kepler they experimented with the SP/DP balance a bit (1/3), in Maxwell they were pushed by 28 nm, in Pascal they returned to the DP = 1/2 SP throughput.

Also, FP16 is not completely separate silicone, AFAIK FP16 instructions are dual-issued on the SP hardware.

> NVIDIA came out with the Titan and Titan Black with baller FP16 performance and no one seem to care

Source? AFAIK earlier FP16 was only supported natively by textures and by some conversion instructions. First chip was the GM20X/Tegra X1 with native FP16.

> As for Google and Baidu while they are huge I'm not sure how "important" they are, Google can pretty much design their own hardware at this point,

Some hardware that allows large benefits, but definitely not all hardware. They're happily relying heavily on GPUs, are planning to pick up Power8+NVlink, etc. Designing chips is expensive, especially if there isn't a huge market to pay for it.

> and as you mentioned they don't really use the software ecosystem that much as they can write everything from scratch even the driver if need be.

Note that they "only" wrote the fronted and IR optimizer, the code generator is NVIDIA's NVPTX [3]!

To wrap up because this is getting long, to your last points I'd say the big DL players are very important for NVIDIA because they are the trendsetters leading in many aspects of AI/DNN research with OSS toolkits for GPUs, and they'd be silly to not use GPUs in-house for their own needs which they're happy to talk about at various conferences and trade shows (e.g. both keynotes GTC 2015 [3] [4]).

[1] http://arstechnica.com/information-technology/2015/12/facebo... [2] http://llvm.org/devmtg/2015-10/slides/Wu-OptimizingLLVMforGP... [3] http://www.ustream.tv/recorded/60071572 [4] http://www.ustream.tv/recorded/60071572

Re: Radeon Instinct – Optimized Machine and Deep Learning

#80
post #55
post #33

Earlier quoted context omitted.

FYI, AMD's people are helping TF to add OpenCL support. https://github.com/tensorflow/tensorflow/issues/22

Contributions from AMD are sorely lacking. It's mostly people from https://www.codeplay.com/ and https://github.com/hughperkins/

I noticed Hugh Perkins has a AMD badge under "Organizations". Is he not working for AMD?

Also one of the earlier participants of the TF OpenCL support conversations https://github.com/gujunli was working for AMD at the time[1].

https://electrek.co/2016/02/26/tesla-machine-learning-expert...

Post reply on HN