Live data from Hacker News

Nvidia Hopper GPU Architecture and H100 Accelerator

anandtech.com

121–130 of 183 posts

Re: Nvidia Hopper GPU Architecture and H100 Accelerator

#121

Earlier quoted context omitted.

Couple points: 1) NVIDIA will likely release a variant of H100 with 2x memory, so we may not even have to wait a generation. They did this for V100-16GB/32GB and A100-40GB/80GB. 2) In a generation or two, the SOTA model architecture will change, so it will be hard to predict the memory reqs... even today, for a fixed train+inference budget, it is much better to train Mixture-Of-Experts (MoE) models, and even NVIDIA a…

Is it difficult/desirable to squeeze/compress an open-sourced 200B parameter model to fit into 40GB? Are these techniques for specific architectures or can they be made generic ?

I think it depends what downstream task you're trying to do... DeepMind tried distilling big language models into smaller ones (think 7B -> 1B) but it didn't work too well... it definitely lost a lot of quality (for general language modeling) relative to the original model.

See the paper here, Figure A28: https://kstatic.googleusercontent.com/files/b068c6c0e64d6f93...

But if your downstream task is simple, like sequence classification, then it may be possible to compress the model without losing much quality.

Re: Nvidia Hopper GPU Architecture and H100 Accelerator

#122
post #67

And may be taking this opportunity to ask, what happen to Nvidia's leak? The hacker hasn't made any more news, and Nvidia hasn't provide an update either.

In the keynote, Jensen made a sly remark about how they themselves could benefit a lot from one of their cyberthreat AI solutions.

Re: Nvidia Hopper GPU Architecture and H100 Accelerator

#123
Seeing the increased bandwidth is super exciting for a lot of business analytics cases we get into for IT/security/fraud/finance teams: imagine correlating across lots of event data from transactions, logs, ... . Every year, it just goes up!

The big welcome surprise for us the secure virtualization. Outside of some limited 24/7 ML teams, we mostly see bursty multi-tenant scenarios for achieving cost-effective utilization. MiG etc static physical partitioning was interesting -- I can imagine cloud providers giving that -- but more dynamic & logical isolation, with more of a focus on namespace isolation, is more relevant to what we see. Once we get into federated learning, and further disintermediations around that, even more cool. Imagine bursting on 0.1-100 GPUs every 30s-20min. Amazing times!

Re: Nvidia Hopper GPU Architecture and H100 Accelerator

#125

80 billion transistors boggles my mind. How many molecules are their per transistor?

It's a crystal so just one molecule for all the transistors. In terms of atoms it's something on the order of the size of a 30 nm cube and with each silicon atom being .2nm in diameter something like 3 million, give or take an order of magnitude or two.

That makes sense. My mistake, I did mean atoms, not molecules. Wolfram alpha estimates 1.35 million Si atoms, so well within 1 order of magnitude.

https://www.wolframalpha.com/input?i=30%5E3+cubic+nanometers...

Re: Nvidia Hopper GPU Architecture and H100 Accelerator

#126
post #44

Earlier quoted context omitted.

Usually the architecture name isn't the only distinguishing feature of the product name, you don't need to remember Intel codenames because a Core 12700 is obviously newer than a Core 11700 Nvidia's accelerators are just called 100 every time so if you don't remember the order of the letters it's not obvious They could have just named them P100, V200, A300 and H400 instead

>you don't need to remember Intel codenames because a Core 12700 is obviously newer than a Core 11700 J3710, 7th Gen J3060, 8th Gen J4205, 8th Gen J4125, 9th Gen i3-5005U, 5th Gen N5095, 10th Gen i7-3770, 3rd Gen 3865U, 7th Gen N3060, 8th Gen

And an AMD 5700U is older than a 5400U as well. A 3400G is older than a 3100X. 3300X isn't really distinctive from 3100X, both are quad-core configurations (but different CCD/cache configurations, which is of course the name doesn't really disclose to the consumer). It happens, naming is a complex topic and there's a lot of dimensions to a product.

In general, complaining about naming is peak bikeshedding for the tech-aware crowd. There are multiple naming schemes, all of them are reasonable, and everyone hates some of them for completely legitimate reasons (but different for every person). And the resulting bikeshedding is exactly as you'd expect with that.

The underlying problem is that products have multiple dimensions of interest - you've got architecture, big vs small core, core count, TDP, clockrate/binning, cache configuration/CCD configuration, graphics configuration, etc. If you sort them by generation, then an older but higher-spec can beat a newer but lower-spec. If you sort by date then refreshes break the scheme. If you split things out into series (m7 vs i7) to express TDP then some people don't like that there's a bunch of different series. If you put them into the same naming scheme then some people don't like that a 5700U is slower than a 5700X. If you try to express all the variables in a single name, you end up with a name like "i7 1185G7" where it's incomprehensible if you don't understand what each of the parts of the name mean.

(as a power user, I personally think the Ice Lake/Tiger Lake naming is the best of the bunch, it expresses everything you need to know: architecture, core count, power, binning, graphics. But then big.LITTLE had to go and mess everything up! And other people still hated it because it was more complex.)

There are certain ones like AMD's 5000 series or the Intel 10th-gen (Comet Lake 10xxxU) that are just really ghastly because they're deliberately trying to mix-and-match to confuse the consumer (to sell older stuff as being new), but in general when people complain about "not understanding all those Lakes and Coves" it's usually just because they aren't interested in the brand/product and don't want to bother learning the names, and they will eagerly rattle off a list of painters or cities that AMD uses as their codenames.

Like, again, to reiterate here, I literally never have seen anyone raise AMD using painter names as being "opaque to the consumer" in the same way that people repeatedly get upset about lakes. And it's the exact same thing. It's people who know the AMD brand and don't know the Intel brand and think that's some kind of a problem with the branding, as opposed to a reflection of their own personal knowledge.

I fully expect that AMD will release 7000 series desktop processors this year or early next year, and exactly 0 people are going to think that a 7600 being newer than a 7702 is confusing in the way that we get all these aggrieved posts about Intel and NVIDIA. Yes, 7600 and 7702 are different product lines, and that's the exact same as your "but i7 3770 and N3060 are different!" example. It's simply not that confusing, it takes less time to learn than to make a single indignant post on social media about it.

Similarly, the NVIDIA practice of using inventors/compsci people is not particularly confusing either. Basically the same as AMD with the painters/cities.

It's just not that interesting, and it's not worth all the bikeshedding that gets devoted to it.

Anyway, your example is all messed up though. J3710 and J3060 are both the same gen (Braswell), launched at the same time (Q1 2016), that example is entirely wrong. J4125 vs J4205 is an older but higher specced processor vs a newer but lower spec, it's a 8th gen Pentium vs a 9th gen Celeron, like a 3100X vs a 2700X (zomg 3100X is bigger number but actually slower!). And the J4125 and J4205 are refreshes of the same architecture with legitimately very similar performance classes. i3 and Atom or i7 and Atom are completely different product lines and the naming is not similar at all there, apart from both having 3s as their first number (not even first character, that is different too, just happen to share the first number somewhere in the name).

Again, like with the Tiger Lake 11xxGxx naming, the characters and positions in the name have meaning. You can come up with better examples than that even within the Intel lineup. Just literally picking 3770 and J3060 as being "similar" because they both have 3s in them.

The one I would legitimately agree on is that the Atom lineup is kind of a mess. Braswell, Apollo Lake, Gemini Lake, and Gemini Lake Refresh are all crammed into the "3000/4000" series space, and there is no "generational number" in that scheme either. Braswell is all 3000 series and Gemini Lake/Gemini Lake Refresh is all 4000 series but you've got Apollo Lake sitting in the middle with both 3000 and 4000 series chips. And a J3455 (Apollo Lake 1.5 GHz) is legitimately a better (or at least equal) processor to a J3710 (Braswell 1.6 GHz). Like 5700U vs 5800U, there are some legitimate architectural differences behind hidden behind an opaque number there (and on the Intel it's graphics - Gemini Lake/Gemini Lake Refresh have a much better video block).

(And that's the problem with "performance rating" approaches, even if a 3710 and a 3455 are similar in performance there's still other differences between them. Also, PR naming instantly turns into gamesmanship - what benchmark, what conditions, what TDP, what level of threading? Is an Intel 37000 the same as an AMD 37000?)

Re: Nvidia Hopper GPU Architecture and H100 Accelerator

#127

Off topic but I can't stand when corporations use actual people's names for their marketing who never gave them the permission to do so. For something like Shakespeare or Cicero I'm OK with it but Grace Hopper was alive in my lifetime, and even Tesla feels a little weird. What gives you the right to use that person's reputation to shill your product?

what gives you the right to own a name ever? especially once your dead?

Re: Nvidia Hopper GPU Architecture and H100 Accelerator

#128

Earlier quoted context omitted.

We have needed wide use of NVlink or something like it for a long time now......heres to hoping mobo manufacturers actually widely implement it!

The open standard version of NVLink is CXL. They're available in latest gen CPUs.

Interesting - I did not know that. Don't we also need motherboard manufacturers though to more widely implement the hardware required? It has been awhile since I have read about NVlink to be fair

Re: Nvidia Hopper GPU Architecture and H100 Accelerator

#129

Earlier quoted context omitted.

> Well the addition is done in FP32 Addition is done in 10-bit mantissa. So maybe TF19 might be the better name, since its a 19-bit format (slightly more than 16-bit BFloats). Really, its a BFloat with a 10-bit mantissa instead of a 7-bit mantissa. 10-bit mantissa matches FP16, while the 8-bit exponent matches FP32. So TF19 probably would have been the best name, but NVidia like marketing so they call it TF32 instead…

It's a 32-bit format in memory and the additions are done with 32-bits.

I admit that I don't have the hardware to test your claims. But pretty much all the whitepapers I can find on TF32 explicitly state the 10-bit mantissa, suggesting that this is at best, a 19-bit format. 1-bit sign + 8-bit exponent + 10-bit mantissa.

Yes, the system will read/write the 32-bit value to RAM. But if there's only 10-bits of mantissa in the circuits, you're only going to get 10-bits of precision (best case). The 10-bit mantissa makes sense because these systems have FP16 circuits (1 + 5-bit exponent + 10-bit mantissa) and BFloat16 circuits (1 sign + 8-bit exponent + 7-bit mantissa). So the 8-bit exponent circuit + 10-bit mantissa circuit exists physically on those NVidia cores.

-------

But the 'Tensor Cores' do not support 32-bit (aka: 23-bit mantissa) or higher.

Post reply on HN