Posts like this underscore why the smart money is betting on Google as the long term AI winner. Meta, Microsoft, OpenAI, etc. are trying to address problems with consumer video cards and spending billions to try and out bid each other to win Nvidia's favor - while Google is on their 6th generation of custom silicon. Literally the only thing that can stop Google now is the fact they keep bringing Microsoft and Oracle…
I think it's likely Nvidia's GPU's, many of which are $50,000+ for a single unit, far surpass Google's custom silicon otherwise why wouldn't Google be selling shovels like Nvidia? If Google had a better chip, or even a chip that was close, they would sell it to anyone and everyone. From a quick search I can see Google's custom chips are 15x to 30x slower to train AI compared to Nvidia's current latest gen AI specific…
How Meta trains large language models at scale
41–50 of 213 posts
Re: How Meta trains large language models at scale
#42Earlier quoted context omitted.
Except Microsoft is making their own chips as well? https://www.theverge.com/2023/11/15/23960345/microsoft-cpu-g...
and so is Meta: https://ai.meta.com/blog/next-generation-meta-training-infer...
Re: How Meta trains large language models at scale
#43Earlier quoted context omitted.
Exactly and they are still about 1/18ths as good at training llms as a H100. Maybe they are less than 1/18ths the cost, so google technically have a marginally better unit cost but i doubt it when you consider the R&D cost. They are less bad at inference, but still much worse than even an A100.
Also energy cost. 18 chips vs 1, it's probably costing a lot more to run 18
Despite various details I don't think that this is an area where Facebook is very different from Google. Both have terrifying amounts of datacenter to play with. Both have long experience making reliable products out of unreliable subsystems. Both have innovative orchestration and storage stacks. Meta hasn't published much or anything about things like reconfigurable optical switches, but that doesn't mean they don't have such a thing.
Re: How Meta trains large language models at scale
#44Random q, I wonder if gloo is used in these systems? https://github.com/facebookincubator/gloo RDMA and GPUDirect capable. Coordinates over MPI or (hi)redia.
Edit: I had a brain freeze or something... gloo is not CPU only but for whatever reason I don't see it outside of CPU-comms
Re: How Meta trains large language models at scale
#45Re: How Meta trains large language models at scale
#46Earlier quoted context omitted.
and so is Meta: https://ai.meta.com/blog/next-generation-meta-training-infer...
Ever since Apple did it everyone has leaped on board. Let's see how things pan out for everyone...
Re: How Meta trains large language models at scale
#47Posts like this underscore why the smart money is betting on Google as the long term AI winner. Meta, Microsoft, OpenAI, etc. are trying to address problems with consumer video cards and spending billions to try and out bid each other to win Nvidia's favor - while Google is on their 6th generation of custom silicon. Literally the only thing that can stop Google now is the fact they keep bringing Microsoft and Oracle…
Is no one else working on custom silicon?
Re: How Meta trains large language models at scale
#48Posts like this underscore why the smart money is betting on Google as the long term AI winner. Meta, Microsoft, OpenAI, etc. are trying to address problems with consumer video cards and spending billions to try and out bid each other to win Nvidia's favor - while Google is on their 6th generation of custom silicon. Literally the only thing that can stop Google now is the fact they keep bringing Microsoft and Oracle…
Is no one else working on custom silicon?
Apple, AWS, Google, Meta, Microsoft all have custom AI-centric silicon.
Re: How Meta trains large language models at scale
#49interesting that their domain is still engineering.fb.com
Re: How Meta trains large language models at scale
#50Posts like this underscore why the smart money is betting on Google as the long term AI winner. Meta, Microsoft, OpenAI, etc. are trying to address problems with consumer video cards and spending billions to try and out bid each other to win Nvidia's favor - while Google is on their 6th generation of custom silicon. Literally the only thing that can stop Google now is the fact they keep bringing Microsoft and Oracle…