Live data from Hacker News

Nvidia CEO Introduces Nvidia Ampere Architecture, Nvidia A100 GPU

blogs.nvidia.com

211–220 of 347 posts

Re: Nvidia CEO Introduces Nvidia Ampere Architecture, Nvidia A100 GPU

#211
post #132

Earlier quoted context omitted.

A fun quick of GPU pricing economics is that on Google Cloud platform, the relatively recent T4 (on Turing and has FP16 support) is cheaper than the ancient K80s. https://cloud.google.com/compute/gpus-pricing

That's not a great comparison at all though. The K80 is a general purpose chip while the T4 is explicitly marketed as an inference chip. The K80 has more ram (super important for batch sizes during training), can access that ram faster (480 GB/s versus 320 GB/s), and is an overall more powerful chip than the T4 is.

Those metrics are for both GPUs on a board; the GCP K80 only uses one GPU, so those performance metrics in theory would be halved (notably, the GCP K80 has 12 GB VRAM vs. T4's 16 GB VRAM), and it's still more expensive than a T4.

Re: Nvidia CEO Introduces Nvidia Ampere Architecture, Nvidia A100 GPU

#212
post #8
post #3

Oof. Double-precision is only 2.5x better, which is less impressive than 20x for float. I still haven't found anything more cost-effective for double precision data processing (cost + dev time) than a rack full of used Xeons...

The double-precision number probably best represents the generational improvement. The 20x 32-bit floating point improvement is probably achieved by comparing doing full FP32 calculations on the previous generation vs doing TF32 calculations on Ampere. This would not be an apple-to-apple comparison as the TF32 result is less precise. That said, it is probably not terribly important for deep learning at least, given t…

Wow. TF32 is only 19 bits. Thats some dubious marketing.

Re: Nvidia CEO Introduces Nvidia Ampere Architecture, Nvidia A100 GPU

#213
post #135
post #120

Earlier quoted context omitted.

I don't underestimate the complexity. But I do claim that the complexity can and should be hidden behind programming language constructs. I've worked both on the design of MIMD hardware, back when I was a graduate student, and on programming languages. These aren't easy problems, but they are solvable. The reason for openness isn't abstract. I don't think NVidia will solve these problems alone. NVidia can make really…

I'm not making a claim about the necessity of experimentation. I spent years (and working a paid job) doing programming language work, and also design hardware these days in my spare time, so I'm not against that. I'm specifically addressing the claim that "GPU databases haven't taken off because of lack of open source CUDA" or whatnot. Database tech is one of the most R&D heavy engineering subfields, almost all majo…

> There is also the problem of needing huge amounts of capital, where most of this work can only be done by exceedingly well funded groups with deep ties to hardware divisions in question. The future of hardware innovation comes from billion dollar companies, because only they can sustain it, not plucky engineers. Sure, for us, CUDA being open source would be awesome. But you don't really need open source drivers when you're working directly with the vendor on your requirements and you pay them millions for support and you just use Linux for everything.

I think the exact same argument could be made for mainframes and microcomputers before we standardized on x86. RISC architectures were cheaper and faster in the eighties and nineties than CISC, but x86 cleaned up because it was standard and had an ecosystem. NVidia is limiting its ecosystem to everyone who needs HPC, where the ecosystem should be everyone (no qualifier). All computers could benefit from a massively parallel MIMD co-processor.

> But what they also understand is that their software stack is a differentiator for them, because it actually works (the competitors don't) and it makes them money to keep it that way.

And I think Symbian made the same argument before being steamrolled by iOS and Android. And I've seen the same argument made by business folks at several businesses I've worked at.

By the way, "open" doesn't mean it's not okay to keep some pieces proprietary. NVidia can keep their differentiator by keeping key algorithms proprietary, while making the architecture open, and developing a common set of cross-platform APIs to target that architecture. For example, a cell phone maker can open source most of their OS, but keep pieces like the fancy ML integrated into their photography app (and other similar pieces) proprietary.

> Nvidia fully understands that maybe some nebulous benefit might come to them by open sourcing things, maybe years down the line.

I think you hit the nail on the head here. The benefits of open feel nebulous; it's a long-tail effect and difficult to quantify. It also takes time. On the other hand, the benefits of proprietary are short-term and easy to quantify. Wrong business decisions get made all the time. Indeed, bad business decisions sometimes get made where everyone can tell it's the wrong decision -- it's just org structures are set up to make those decisions. I think this isn't me claiming to be brilliant or smarter than NVidia so much as NVidia failing in the same exact way many organizations fail, by the design of the org structures.

> They understand plucky researchers can do amazing things, sometimes.

It's actually not just about amazing things. It's about a long tail of dumb stuff too. My phone has a few apps better than Google could build. It has dozens of apps Google chose not to build. Most of the stuff I want to do isn't big enough to ever show up on NVidia's radar, but there are a lot of people like me. Symbian didn't make a piano tuner app. It's not hard to make one. I have one, though.

Of course, there are brilliant pieces too. I have some VR/AR apps on my phone which Google would need to invest a lot of capital to make.

> EDIT: I'll also say that if this changes from their "major open source announcement" they were going to do at GTC, I'll eat my hat. I'm not expecting much from Nvidia in terms of open source, but I'd happily be proven wrong. But broadly I think my general point stands, which is that thinking about it from the POV of "open source drivers are the limitation" isn't really the right way to think about it.

I'm not holding my breath for NVidia to change. But I do hope at some point, we'll see a nice, open MIMD architecture which gives me that nice 10-100x speedup for parallel workloads. I actually couldn't care less about whether that speed-up is 50x or 100x (which is where NVidia's deep R&D advantage lies). That matters for bitcoin mining or deep learning. For the long tail I'm talking about, the baseline speedup is plenty good enough. The cleverness doesn't come from pushing extra CPU cycles out, but in APIs, ecosystem building, openness, standardization, etc. That stuff is a different kind of hard.

Re: Nvidia CEO Introduces Nvidia Ampere Architecture, Nvidia A100 GPU

#214
post #140
post #120

Earlier quoted context omitted.

I don't underestimate the complexity. But I do claim that the complexity can and should be hidden behind programming language constructs. I've worked both on the design of MIMD hardware, back when I was a graduate student, and on programming languages. These aren't easy problems, but they are solvable. The reason for openness isn't abstract. I don't think NVidia will solve these problems alone. NVidia can make really…

Except there is a community, a CUDA community and from GTC sessions, a very big one. Ironically this walled garden as you put it, has produced more programming languages and tooling for GPGPU programming than the open conglomerate design by committee from Khronos has been able to achieve together against a single company, which kept pushing their C mantra until it was too late.

Yes. To have openness work, you need to execute well too.

Re: Nvidia CEO Introduces Nvidia Ampere Architecture, Nvidia A100 GPU

#215
post #158

Earlier quoted context omitted.

Tesla is famous in the US for inventing polyphase AC and induction motors, but this is really one of these stories were a bunch of people invented the same thing very closely to each other due to a precipitating reaching of understanding. Note that Tesla's designs were IIRC two-phase which is largely inferior to three-phase. The push for three-phase and associated designs and inventions (three phase transformers on a…

> but this is really one of these stories were a bunch of people invented the same thing very closely to each other due to a precipitating reaching of understanding. This is by far the dominant case of invention. Truly independent work is incredibly rare.

Indeed. I found it quite enlightening to go through this list: https://en.wikipedia.org/wiki/List_of_multiple_discoveries

Re: Nvidia CEO Introduces Nvidia Ampere Architecture, Nvidia A100 GPU

#216
post #93
post #31

Earlier quoted context omitted.

It's irrelevant to researchers. Research operates on rapid cycles: prototype, publish, move on. It does impact businesses. It doesn't prevent adoption for e.g. deep learning, but I haven't seen e.g. GPU-based databases reach broad adoption, or many other places where MIMD/SIMD would reduce costs or improve performance. Using classical hardware is clearly cheaper than the business risk and engineering time of relying…

NVidia is not to blame if the competition is stuck using C, printf debugging for computing shaders, cannot make their minds about which bytecode to support for heterogenous GPGPU programming. The situation is so bad that OpenCL 1.2 got promoted to OpenCL 3.0 and SYSCL is now backend independent, while hip only works on Linux. As for Python, guess who is on the forefront of GPU Programming with Python, https://www.nvi…

It's not about blame. It's about getting to an ecosystem where GPGPU is used for things beyond deep learning, bitcoin mining, video encoding, and similar niche applications to one where I can fluidly use MIMD to speed up my JavaScript and Python with first-class language constructs to support that.

If that happens:

1) We'll get back on some kind of curve where computer performance starts increasing again.

2) The GPU will become more important than the CPU, and the market will explode.

Until that happens, the GPU market will be for gamers, video editors, and machine learning nerds.

I don't much care who does that, or why it hasn't happened.

Re: Nvidia CEO Introduces Nvidia Ampere Architecture, Nvidia A100 GPU

#218
post #152

Earlier quoted context omitted.

> Correction: Nobody will be able to use the AMD hardware (outside of computer graphics) because everybody has been locked-in with CUDA on Nvidia. NVIDA open-sourced their CUDA implementation to the LLVM project 5 years ago, which is why clang can compile CUDA today, and why Intel and PGI have clang forks compiling CUDA to multi-threaded and vectorized x86-64 using OpenMP. That you can't compile CUDA to AMD GPUs isn'…

> Do you work for AMD I do not. And I use NVidia hardware regularly for GPGPU. But I hate fanboyism. > NVIDA open-source their CUDA implementation to the LLVM project 5 years ago Correction: Google developped an internal CUDA implementation for their own need based on LLVM that Nvidia barely supported it for their own need afterwards. Nothing is "stable" nor "branded" in this work.... Consequently, 99% of public Open…

> But I hate fanboyism.

"The only excuse I can see to this attitude is greed" sounds pretty fanboyish to me. :-)

I've never understood why Microsoft, or Adobe, or Autodesk, or Synopsys, or Cadence or any other pure software company is allowed to charge as much as the market will bear for their products, often more per year than Nvidia's hardware, but when a company makes software that runs on dedicated hardware, it's called greed. I don't think it's an exaggeration when I say that, for many laptops with a Microsoft Office 365 license, you pay more over the lifetime of the laptop for the software license than for the hardware itself. And it's definitely true for most workstation software.

When you use Photoshop for your creative work, you lock your design IP to Adobe's Creative Suite. When you use CUDA to create your own compute IP, you lock yourself to Nvidia's hardware.

In both cases, you're going to pay an external party. In both cases, you decide that this money provides enough value to be worth paying for.

Re: Nvidia CEO Introduces Nvidia Ampere Architecture, Nvidia A100 GPU

#219
post #18

I'm confused. Is there any relationship between the recent Ampere Arm64 servers ( https://news.ycombinator.com/item?id=22475036 ) and Nvidia's "Ampere Architecture", or is it just a case of them using the same name?

I don't like that people downvoted you for asking a question. If someone thinks the question is stupid or not doesn't mean that a downvote is warranted. (nor an upvote, answer the question and move on.) To answer though; it's just a coincidence, as you might already know Nvidia uses famous scientists (especially in the field of electricity) as the names of their microarchitectures. * Volta (Alessandro Volta, inventor…

Not just electricity, but basic physics. Don't forget the Kepler architecture, named after the astronomer Johannes Kepler, and the Fermi architecture, named after the nuclear physicist Enrico Fermi (although the same physics is important in semiconductors). Although there;s also an odd exception, which is the Turing architecture.

Re: Nvidia CEO Introduces Nvidia Ampere Architecture, Nvidia A100 GPU

#220

Earlier quoted context omitted.

> Intel keeps getting it right with numerical libraries. They're open. They work well. They work on AMD. What Intel numerical libraries are you thinking of? When I think of Intel numerical libraries, the first that comes to mind is MKL. MKL is neither open-source nor does it work well on AMD without some fragile hacks [0]. [0] https://www.pugetsystems.com/labs/hpc/How-To-Use-MKL-with-AM...

Well, OP didn't say MKL works well on AMD. But you can at least run it on a non-Intel CPU. Compare CUDA.

Not on ARM, or POWER, you can't. Why you'd want to run it on AMD, I don't understand. I don't know what fraction of peak BLIS and OpenBLAS get, but it will be high.
Post reply on HN