Live data from Hacker News

How Meta trains large language models at scale

engineering.fb.com

41–50 of 213 posts

Re: How Meta trains large language models at scale

#41
post #4

Posts like this underscore why the smart money is betting on Google as the long term AI winner. Meta, Microsoft, OpenAI, etc. are trying to address problems with consumer video cards and spending billions to try and out bid each other to win Nvidia's favor - while Google is on their 6th generation of custom silicon. Literally the only thing that can stop Google now is the fact they keep bringing Microsoft and Oracle…

I think it's likely Nvidia's GPU's, many of which are $50,000+ for a single unit, far surpass Google's custom silicon otherwise why wouldn't Google be selling shovels like Nvidia? If Google had a better chip, or even a chip that was close, they would sell it to anyone and everyone. From a quick search I can see Google's custom chips are 15x to 30x slower to train AI compared to Nvidia's current latest gen AI specific…

Nvidia has decades of experience selling hardware to people with all the pains that entails, support, sales channels, customer acquisition, software, it's something you don't just do overnight, and it does cost money. Google's TPUs get some of their cost efficiency from not supporting COTS use cases and the overhead of selling to people, and the total wall clock time has to also include the total operational costs, which dominate at their size (e.g. if it's 30x slower but 1/50th the TCO then it's a win. I don't know how TPUv5 stacks up against the B200). It's not as simple as "just put it on a shelf and sell it and make a gajillion dollars like nvidia"

Re: How Meta trains large language models at scale

#42
post #30
post #17

Earlier quoted context omitted.

Except Microsoft is making their own chips as well? https://www.theverge.com/2023/11/15/23960345/microsoft-cpu-g...

and so is Meta: https://ai.meta.com/blog/next-generation-meta-training-infer...

Ever since Apple did it everyone has leaped on board. Let's see how things pan out for everyone...

Re: How Meta trains large language models at scale

#43
post #39
post #28

Earlier quoted context omitted.

Exactly and they are still about 1/18ths as good at training llms as a H100. Maybe they are less than 1/18ths the cost, so google technically have a marginally better unit cost but i doubt it when you consider the R&D cost. They are less bad at inference, but still much worse than even an A100.

Also energy cost. 18 chips vs 1, it's probably costing a lot more to run 18

Google claims the opposite in "TPU v4: An Optically Reconfigurable Supercomputer for Machine Learning with Hardware Support for Embeddings " https://arxiv.org/abs/2304.01433

Despite various details I don't think that this is an area where Facebook is very different from Google. Both have terrifying amounts of datacenter to play with. Both have long experience making reliable products out of unreliable subsystems. Both have innovative orchestration and storage stacks. Meta hasn't published much or anything about things like reconfigurable optical switches, but that doesn't mean they don't have such a thing.

Re: How Meta trains large language models at scale

#44

Random q, I wonder if gloo is used in these systems? https://github.com/facebookincubator/gloo RDMA and GPUDirect capable. Coordinates over MPI or (hi)redia.

̶I̶I̶R̶C̶ ̶g̶l̶o̶o̶ ̶i̶s̶ ̶C̶P̶U̶ ̶t̶e̶n̶s̶o̶r̶s̶ ̶o̶n̶l̶y̶ ̶s̶o̶ ̶l̶i̶k̶e̶l̶y̶ ̶n̶o̶t̶

Edit: I had a brain freeze or something... gloo is not CPU only but for whatever reason I don't see it outside of CPU-comms

Re: How Meta trains large language models at scale

#46
post #42
post #30

Earlier quoted context omitted.

and so is Meta: https://ai.meta.com/blog/next-generation-meta-training-infer...

Ever since Apple did it everyone has leaped on board. Let's see how things pan out for everyone...

And that's why $ARM is a good buy. Selling swords and steel to all these armies as they go to war.

Re: How Meta trains large language models at scale

#47
post #4

Posts like this underscore why the smart money is betting on Google as the long term AI winner. Meta, Microsoft, OpenAI, etc. are trying to address problems with consumer video cards and spending billions to try and out bid each other to win Nvidia's favor - while Google is on their 6th generation of custom silicon. Literally the only thing that can stop Google now is the fact they keep bringing Microsoft and Oracle…

Is no one else working on custom silicon?

The problem isn’t just developing your own processor. Nvidia has a huge stack of pretty cutting edge technology including a huge stack from mellanox, an enormous OSS tool chain around CUDA, etc, that people seeking to make comparable products have to overcome.

Re: How Meta trains large language models at scale

#48
post #4

Posts like this underscore why the smart money is betting on Google as the long term AI winner. Meta, Microsoft, OpenAI, etc. are trying to address problems with consumer video cards and spending billions to try and out bid each other to win Nvidia's favor - while Google is on their 6th generation of custom silicon. Literally the only thing that can stop Google now is the fact they keep bringing Microsoft and Oracle…

Is no one else working on custom silicon?

Everyone is.

Apple, AWS, Google, Meta, Microsoft all have custom AI-centric silicon.

Re: How Meta trains large language models at scale

#49

interesting that their domain is still engineering.fb.com

I think fb.com is their internal domain and they never really bothered to change it, I think employees used to have a @fb.com e-mail, at least this was true a few years ago, not sure if that has changed.

Re: How Meta trains large language models at scale

#50
post #4

Posts like this underscore why the smart money is betting on Google as the long term AI winner. Meta, Microsoft, OpenAI, etc. are trying to address problems with consumer video cards and spending billions to try and out bid each other to win Nvidia's favor - while Google is on their 6th generation of custom silicon. Literally the only thing that can stop Google now is the fact they keep bringing Microsoft and Oracle…

I would call it "stupid" money. This isn't a commodity business. Value of the final product is orthogonal to amount invested in compute. If Google is 10% slower or its product is 10% worse, it can lose all the value. This is like valuing a software company higher because its devs are using cheap PC desktops instead of Mac.
Post reply on HN