Live data from Hacker News

How Meta trains large language models at scale

engineering.fb.com

51–60 of 213 posts

Re: How Meta trains large language models at scale

#51
post #24

Earlier quoted context omitted.

Do you really think Google’s hardware expertise is better than Nvidia’s? If needed these other companies have the $$$ to buy the best chips money can buy from Nvidia. Better chips than Google could ever produce. If anything, this is why IMO Google will fail.

Yes. Google was building custom HPC hardware 5-8 years before Nvidia decided to expand outside the consumer and "workstation" markets.

Nvidia acquired Mellanox who know far more about custom HPC hardware than Google.

Re: How Meta trains large language models at scale

#52
OK this was a bit funny:

  Top HW failure modes: 
  * GPU falling off the bus
I honestly thought "do they mean GPUs falling off a bus entering the data center" and then realized its actually the connectivity, as they mention in the next line

  GPUs falling off: In this case, GPUs are not detected by the host on PCIe.

Re: How Meta trains large language models at scale

#53
post #4

Posts like this underscore why the smart money is betting on Google as the long term AI winner. Meta, Microsoft, OpenAI, etc. are trying to address problems with consumer video cards and spending billions to try and out bid each other to win Nvidia's favor - while Google is on their 6th generation of custom silicon. Literally the only thing that can stop Google now is the fact they keep bringing Microsoft and Oracle…

I wish I had your faith in Google’s ability to refrain from kneecapping their own perfectly fine product.

Re: How Meta trains large language models at scale

#54
These seem classic challenges with running distributed systems loads that are not specific to training LLMs.

Anyone of the super computers listed here https://en.wikipedia.org/wiki/TOP500 suffers from the same issues.

Think about it. While the national labs use these systems to model serious stuff -such as climate or nuclear weapons- Meta uses them to train LLMs. What a joke, honestly!

Re: How Meta trains large language models at scale

#57
post #4

Posts like this underscore why the smart money is betting on Google as the long term AI winner. Meta, Microsoft, OpenAI, etc. are trying to address problems with consumer video cards and spending billions to try and out bid each other to win Nvidia's favor - while Google is on their 6th generation of custom silicon. Literally the only thing that can stop Google now is the fact they keep bringing Microsoft and Oracle…

That’s how all huge tech companies become dinosaurs though. Upper management that is already stupidly wealthy (and therefore unmotivated) have the funding and patience to hire geniuses to build incredible machines and then constantly tie their shoelaces together while asking them to sprint. Examples include Microsoft and Oracle as you said, and before them IBM, AT$T, TIBCO, Marvell, Motorola, I could go on for a while…

Re: How Meta trains large language models at scale

#58
post #56

> So we decided to build both: two 24k clusters, one with RoCE and another with InfiniBand. Our intent was to build and learn from the operational experience. I love how they built two completely insane clusters just to learn. That's badass.

More like Mark gave them 100k GPUs, and they are not sure what exactly to do with them..

Re: How Meta trains large language models at scale

#59
post #4

Posts like this underscore why the smart money is betting on Google as the long term AI winner. Meta, Microsoft, OpenAI, etc. are trying to address problems with consumer video cards and spending billions to try and out bid each other to win Nvidia's favor - while Google is on their 6th generation of custom silicon. Literally the only thing that can stop Google now is the fact they keep bringing Microsoft and Oracle…

I think it's likely Nvidia's GPU's, many of which are $50,000+ for a single unit, far surpass Google's custom silicon otherwise why wouldn't Google be selling shovels like Nvidia? If Google had a better chip, or even a chip that was close, they would sell it to anyone and everyone. From a quick search I can see Google's custom chips are 15x to 30x slower to train AI compared to Nvidia's current latest gen AI specific…

We have almost 400 H100's sitting idle. I wonder how many other companies are buying millions of dollars worth of these chips with the hopes of them being used, but aren't being utilized?

Re: How Meta trains large language models at scale

#60
post #4

Posts like this underscore why the smart money is betting on Google as the long term AI winner. Meta, Microsoft, OpenAI, etc. are trying to address problems with consumer video cards and spending billions to try and out bid each other to win Nvidia's favor - while Google is on their 6th generation of custom silicon. Literally the only thing that can stop Google now is the fact they keep bringing Microsoft and Oracle…

None of these companies are using consumer video cards. https://www.nvidia.com/en-us/data-center/h200/
Post reply on HN