> Not sure if Maxwell came out of the blue, Maxwell 1 was designed for mobile, embedded and tesla, Maxwell 2 came out most likely because Pascal at large was delayed.
First off not sure what you are referring to by "Maxwell 1" and "Maxwell 2"; there's GM204, GM206, GM200 all very similar, and GM20B the slight outlier.
Maxwell was an arch tweak on Kepler which, due to everyone but Intel stuck at 28 nm, had to cut down on all but the gaming-essential stuff to deliver what the consumers expected (>1.5x gen-boost). For that reason, and because HPC has traditionally been thought to need DP, IIRC no public roadmap mentioned Maxwell Tesla at all. Instead, the GK210 (K80) was the attempt to be the bridge-gap chip for HPC until Pascal; it was released just a few months before the big GM200 consumer chips came out. That is until fall '15 when the late arriving M40 and M4 were pitched as "Deep learning accelerators" (though GRID/virtualization version were released a little earlier in August). Quite obvious naming, plus no sane HPC shop would buy a crippled (no DP) chip > As for the half precision, it's pretty much the same thing NVIDIA been doing since Kepler dumping FP64 and FP16, especially FP16 due to the silicon costs.
Not really. In Kepler they experimented with the SP/DP balance a bit (1/3), in Maxwell they were pushed by 28 nm, in Pascal they returned to the DP = 1/2 SP throughput.
Also, FP16 is not completely separate silicone, AFAIK FP16 instructions are dual-issued on the SP hardware.
> NVIDIA came out with the Titan and Titan Black with baller FP16 performance and no one seem to care
Source? AFAIK earlier FP16 was only supported natively by textures and by some conversion instructions. First chip was the GM20X/Tegra X1 with native FP16.
> As for Google and Baidu while they are huge I'm not sure how "important" they are, Google can pretty much design their own hardware at this point,
Some hardware that allows large benefits, but definitely not all hardware. They're happily relying heavily on GPUs, are planning to pick up Power8+NVlink, etc. Designing chips is expensive, especially if there isn't a huge market to pay for it.
> and as you mentioned they don't really use the software ecosystem that much as they can write everything from scratch even the driver if need be.
Note that they "only" wrote the fronted and IR optimizer, the code generator is NVIDIA's NVPTX [3]!
To wrap up because this is getting long, to your last points I'd say the big DL players are very important for NVIDIA because they are the trendsetters leading in many aspects of AI/DNN research with OSS toolkits for GPUs, and they'd be silly to not use GPUs in-house for their own needs which they're happy to talk about at various conferences and trade shows (e.g. both keynotes GTC 2015 [3] [4]).
[1] http://arstechnica.com/information-technology/2015/12/facebo...
[2] http://llvm.org/devmtg/2015-10/slides/Wu-OptimizingLLVMforGP...
[3] http://www.ustream.tv/recorded/60071572
[4] http://www.ustream.tv/recorded/60071572