What's particularly interesting here is that the Fiji card they propose is a very different beast than any of the NVIDIA offerings. The MI8 card's HBM has a great power and performance advantage (512 GB/s peak bandwidth) even if it's on 28 nm. NVIDIA has nothing that has even remotely comparable bandwidth in this price/perf/TDP regime. None of the NVIDIA GP10[24] Teslas have GDDR5X -- not to surprising given that it…
Those MIOpen benchmarks are a bit dubious, since MIOpen is AMDs own deep learning framework. It's unlikely that code written by AMD is optimal for the Nvidia hardware. To be realistic you need to compare AMD hardware running MIOpen to NV hardware running a framework backed by cuDNN.
Radeon Instinct – Optimized Machine and Deep Learning
81–88 of 88 posts
Re: Radeon Instinct – Optimized Machine and Deep Learning
#82Earlier quoted context omitted.
Not sure if Maxwell came out of the blue, Maxwell 1 was designed for mobile, embedded and tesla, Maxwell 2 came out most likely because Pascal at large was delayed. As for the half precision, it's pretty much the same thing NVIDIA been doing since Kepler dumping FP64 and FP16, especially FP16 due to the silicon costs. NVIDIA came out with the Titan and Titan Black with baller FP16 performance and no one seem to care,…
> Not sure if Maxwell came out of the blue, Maxwell 1 was designed for mobile, embedded and tesla, Maxwell 2 came out most likely because Pascal at large was delayed. First off not sure what you are referring to by "Maxwell 1" and "Maxwell 2"; there's GM204, GM206, GM200 all very similar, and GM20B the slight outlier. Maxwell was an arch tweak on Kepler which, due to everyone but Intel stuck at 28 nm, had to cut down…
GM1XX is Maxwell 1st Gen, which was the 750ti, the Gefore 800M series, Tegra K1 and Tesla M10, GM2XX is Maxwell 2nd Gen.
>Not really. In Kepler they experimented with the SP/DP balance a bit (1/3), in Maxwell they were pushed by 28 nm, in Pascal they returned to the DP = 1/2 SP throughput
Maxwell wasn't 1/3 it was 1/32 (no this isn't a typo) ;) this is the same (or even worse IIRC it's 1/64 now) with Pascal with the exclusion of the GP100, the Pascal Titan and Quadro cards are 1/32 or 1/64.
As I said NVIDIA keeps this for very limited silicon if you buy a desktop/workstation GPU even with Pascal don't expect DP/HP performance.
As for the big players, yes they are trend setters but they are also can easily turn into "blackmailers" if you have a single client which buys half your GPU's they dictate the terms, companies went under because of contracts that take too much of their delivery pipeline.
Re: Radeon Instinct – Optimized Machine and Deep Learning
#83Earlier quoted context omitted.
Contributions from AMD are sorely lacking. It's mostly people from https://www.codeplay.com/ and https://github.com/hughperkins/
I noticed Hugh Perkins has a AMD badge under "Organizations". Is he not working for AMD? Also one of the earlier participants of the TF OpenCL support conversations https://github.com/gujunli was working for AMD at the time[1]. https://electrek.co/2016/02/26/tesla-machine-learning-expert...
AMD could really invest more in optimising deep learning libraries for their hardware if they want to be relevant
Re: Radeon Instinct – Optimized Machine and Deep Learning
#84Earlier quoted context omitted.
I noticed Hugh Perkins has a AMD badge under "Organizations". Is he not working for AMD? Also one of the earlier participants of the TF OpenCL support conversations https://github.com/gujunli was working for AMD at the time[1]. https://electrek.co/2016/02/26/tesla-machine-learning-expert...
No, Hugh works at asapp.com. AMD could really invest more in optimising deep learning libraries for their hardware if they want to be relevant
> AMD could really invest more in optimising deep learning libraries for their hardware if they want to be relevant
I am optimistic about AMD's involvement in AI. Just do a search of recent news on here, tons of exciting devleopment from AMD, especially Radeon Instinct. And here is an interesting comment about the comparison between Nvidia and AMD[1].
Re: Radeon Instinct – Optimized Machine and Deep Learning
#85Earlier quoted context omitted.
> Not sure if Maxwell came out of the blue, Maxwell 1 was designed for mobile, embedded and tesla, Maxwell 2 came out most likely because Pascal at large was delayed. First off not sure what you are referring to by "Maxwell 1" and "Maxwell 2"; there's GM204, GM206, GM200 all very similar, and GM20B the slight outlier. Maxwell was an arch tweak on Kepler which, due to everyone but Intel stuck at 28 nm, had to cut down…
>First off not sure what you are referring to by "Maxwell 1" and "Maxwell 2"; there's GM204, GM206, GM200 all very similar, and GM20B the slight outlier. GM1XX is Maxwell 1st Gen, which was the 750ti, the Gefore 800M series, Tegra K1 and Tesla M10, GM2XX is Maxwell 2nd Gen. >Not really. In Kepler they experimented with the SP/DP balance a bit (1/3), in Maxwell they were pushed by 28 nm, in Pascal they returned to the…
My bad, forgot about 5.0 devices being called GM1xx. Still, AFAIR there was little practical difference (in particular instruction set) between all but the compute capability 5.3 Tegra.
> Maxwell wasn't 1/3 it was 1/32 (no this isn't a typo) ;) this is the same (or even worse IIRC it's 1/64 now) with Pascal with the exclusion of the GP100, the Pascal Titan and Quadro cards are 1/32 or 1/64.
I did not say Maxwell had 1/3 DP flop rate, I said Kepler had that, but Maxwell was pushed by the 28 nm process (so they had to get rid of all DP ALUs).
> As I said NVIDIA keeps this for very limited silicon if you buy a desktop/workstation GPU even with Pascal don't expect DP/HP performance.
Never disagreed, but they can only do that because their professional compute division has grown big enough that it it worth designing different silicon; that shift happened in fact after Kepler, GK110 was more or less the same on GeForce and Tesla (e.g. 780 Ti and K40). They're also trying to find a good way to do market segmentation and DP ALU die area is an obvious candidate to play with. Still, no HP on GP102/104 is a curious thing, especially as the Tesla P4/P40 that oficially target ML/DL would really need it. My bet is they won't "forget" HP on lower-end Volta; not all consumer chips will support it, but unless they'll design different dies for GV102/104 (or whatever they'll call the non-DP Tesla uarch), which I doubt, some desktop parts will also support it.
Re: Radeon Instinct – Optimized Machine and Deep Learning
#86Earlier quoted context omitted.
I remember that a few years ago AMD had the only sensible solution for virtualizing GPUs, and you could make a bunch of them work together as a single unit without much trouble. But I didn't have a Radeon card so I never got to try it. Don't know what happened with that, it's true that they really lagged behind NVIDIA.
AMD has a new virtualization solution which provides SR-IOV hardware based partitioning on their AMD FirePro S7150 or S7150 x2 gpu cards. http://www.amd.com/en-us/solutions/professional/virtualizati... This video explains what you get https://www.youtube.com/watch?v=tKBthlKTtvQ
Re: Radeon Instinct – Optimized Machine and Deep Learning
#87Earlier quoted context omitted.
>First off not sure what you are referring to by "Maxwell 1" and "Maxwell 2"; there's GM204, GM206, GM200 all very similar, and GM20B the slight outlier. GM1XX is Maxwell 1st Gen, which was the 750ti, the Gefore 800M series, Tegra K1 and Tesla M10, GM2XX is Maxwell 2nd Gen. >Not really. In Kepler they experimented with the SP/DP balance a bit (1/3), in Maxwell they were pushed by 28 nm, in Pascal they returned to the…
> GM1XX is Maxwell 1st Gen, which was the 750ti, the Gefore 800M series, Tegra K1 and Tesla M10, GM2XX is Maxwell 2nd Gen. My bad, forgot about 5.0 devices being called GM1xx. Still, AFAIR there was little practical difference (in particular instruction set) between all but the compute capability 5.3 Tegra. > Maxwell wasn't 1/3 it was 1/32 (no this isn't a typo) ;) this is the same (or even worse IIRC it's 1/64 now)…
As for Kepler well it had 1:3 but only on few chips, most kepler based GPU's had 1:24 DP performance, the 780ti had 1:24 while the Kepler Titan had 1:3, the Titan Black then had 1:24 again and again no one seemed to care...
Re: Radeon Instinct – Optimized Machine and Deep Learning
#88Earlier quoted context omitted.
There are several well known (certainly to AMD) strategies and tactics for dealing with this situation. One example: AMD can support CUDA while focusing on price/performance at the low end (initiating an Innovators Dilemma for NVIDIA), and if successful then solidify AMD's position, for example by starting a standards process for a successor to CUDA. NVIDIA, of course, has various countering moves available such as I…
> if successful then solidify AMD's position, for example by starting a standards process for a successor to CUDA. I thought OpenCL was the standard in this space.
XHTML 2.0 comes to mind.
Tactically, creating a backward-compatible successor to CUDA (a superset, essentially) would have the advantage (for AMD, that is) that NVIDIA can't decide to un-implement the existing CUDA support in their products in order to to spike AMD's efforts.
Then again, there is a lot of IP sloshing around in this space, so the specific tactic used by AMD may have to be something sneakier and subtler.
Anyway, this is all just one of the hypothetical ways AMD could try to unseat NVIDIA, there are others.