Earlier quoted context omitted.
Do you really think Google’s hardware expertise is better than Nvidia’s? If needed these other companies have the $$$ to buy the best chips money can buy from Nvidia. Better chips than Google could ever produce. If anything, this is why IMO Google will fail.
Yes. Google was building custom HPC hardware 5-8 years before Nvidia decided to expand outside the consumer and "workstation" markets.
How Meta trains large language models at scale
51–60 of 213 posts
Re: How Meta trains large language models at scale
#52 Top HW failure modes:
* GPU falling off the bus
I honestly thought "do they mean GPUs falling off a bus entering the data center" and then realized its actually the connectivity, as they mention in the next line GPUs falling off: In this case, GPUs are not detected by the host on PCIe.Re: How Meta trains large language models at scale
#53Posts like this underscore why the smart money is betting on Google as the long term AI winner. Meta, Microsoft, OpenAI, etc. are trying to address problems with consumer video cards and spending billions to try and out bid each other to win Nvidia's favor - while Google is on their 6th generation of custom silicon. Literally the only thing that can stop Google now is the fact they keep bringing Microsoft and Oracle…
Re: How Meta trains large language models at scale
#54Anyone of the super computers listed here https://en.wikipedia.org/wiki/TOP500 suffers from the same issues.
Think about it. While the national labs use these systems to model serious stuff -such as climate or nuclear weapons- Meta uses them to train LLMs. What a joke, honestly!
Re: How Meta trains large language models at scale
#55Re: How Meta trains large language models at scale
#56I love how they built two completely insane clusters just to learn. That's badass.
Re: How Meta trains large language models at scale
#57Posts like this underscore why the smart money is betting on Google as the long term AI winner. Meta, Microsoft, OpenAI, etc. are trying to address problems with consumer video cards and spending billions to try and out bid each other to win Nvidia's favor - while Google is on their 6th generation of custom silicon. Literally the only thing that can stop Google now is the fact they keep bringing Microsoft and Oracle…
Re: How Meta trains large language models at scale
#58> So we decided to build both: two 24k clusters, one with RoCE and another with InfiniBand. Our intent was to build and learn from the operational experience. I love how they built two completely insane clusters just to learn. That's badass.
Re: How Meta trains large language models at scale
#59Posts like this underscore why the smart money is betting on Google as the long term AI winner. Meta, Microsoft, OpenAI, etc. are trying to address problems with consumer video cards and spending billions to try and out bid each other to win Nvidia's favor - while Google is on their 6th generation of custom silicon. Literally the only thing that can stop Google now is the fact they keep bringing Microsoft and Oracle…
I think it's likely Nvidia's GPU's, many of which are $50,000+ for a single unit, far surpass Google's custom silicon otherwise why wouldn't Google be selling shovels like Nvidia? If Google had a better chip, or even a chip that was close, they would sell it to anyone and everyone. From a quick search I can see Google's custom chips are 15x to 30x slower to train AI compared to Nvidia's current latest gen AI specific…
Re: How Meta trains large language models at scale
#60Posts like this underscore why the smart money is betting on Google as the long term AI winner. Meta, Microsoft, OpenAI, etc. are trying to address problems with consumer video cards and spending billions to try and out bid each other to win Nvidia's favor - while Google is on their 6th generation of custom silicon. Literally the only thing that can stop Google now is the fact they keep bringing Microsoft and Oracle…