Live data from Hacker News

How Meta trains large language models at scale

engineering.fb.com

31–40 of 213 posts

Re: How Meta trains large language models at scale

#32
post #4

Posts like this underscore why the smart money is betting on Google as the long term AI winner. Meta, Microsoft, OpenAI, etc. are trying to address problems with consumer video cards and spending billions to try and out bid each other to win Nvidia's favor - while Google is on their 6th generation of custom silicon. Literally the only thing that can stop Google now is the fact they keep bringing Microsoft and Oracle…

Much the same way you can have all the best gear and still fail - Google’s primary strength seems to be the DeepMind group. I’m not affiliated with Google, but IMHO the reason they will slowly die is because their engineering culture has taken a backseat due to their broken hiring practices.

Bad hiring practices aren’t exclusive to them, but from all accounts it seems like their internal focus is on optimizing ad revenue over everything else. I could be wrong or misinformed, but it seems to me like they are playing the finite game in the AI space (DeepMind group aside) while FAIR are playing the infinite game.

*meanwhile MSFT are simply trying to buy their way to relevance (e.g. OpenAI investments, etc) and carve out future revenues (Recall) and Jobs-less Apple is building their trademark walled-garden (AppleIntelligence?). Although the use of unified memory in Apple silicon poses some interesting possibilities for enabling the use of sizable models on consumer hardware.

Overall it seems like “big-tech” is by-and-large uninspired and asleep at the wheel save specific teams like those led by Lecun, Hassabis, etc. not sure where that leaves OpenAI now that Karpathy is gone.

Re: How Meta trains large language models at scale

#33
post #9

Earlier quoted context omitted.

[flagged]

Because it's completely irrelevant.

and deceptive if not inaccurate. Meta's Model Cards specifically call out that they were trained on publicly available datasets and NOT any Meta user data.

For example: https://github.com/meta-llama/llama3/blob/main/MODEL_CARD.md

Re: How Meta trains large language models at scale

#35
post #4

Posts like this underscore why the smart money is betting on Google as the long term AI winner. Meta, Microsoft, OpenAI, etc. are trying to address problems with consumer video cards and spending billions to try and out bid each other to win Nvidia's favor - while Google is on their 6th generation of custom silicon. Literally the only thing that can stop Google now is the fact they keep bringing Microsoft and Oracle…

"Consumer video cards"? Meta's not building their clusters out of 3090s.

They're using advanced cards meant for data centers and machine learning -- almost effectively "custom silicon"

Re: How Meta trains large language models at scale

#36
post #16
post #4

Posts like this underscore why the smart money is betting on Google as the long term AI winner. Meta, Microsoft, OpenAI, etc. are trying to address problems with consumer video cards and spending billions to try and out bid each other to win Nvidia's favor - while Google is on their 6th generation of custom silicon. Literally the only thing that can stop Google now is the fact they keep bringing Microsoft and Oracle…

The only thing that can stop Google is Google. Somehow every bet that isn't Search doesn't pan out. And inexplicably, they're working hard to kill Search now. As a shareholder, I hope they succeed. But I am more pessimistic about it than you.

[deleted]

Re: How Meta trains large language models at scale

#38
post #28

Earlier quoted context omitted.

They do sell shovels, you can get Google TPUs on Google Cloud.

Exactly and they are still about 1/18ths as good at training llms as a H100. Maybe they are less than 1/18ths the cost, so google technically have a marginally better unit cost but i doubt it when you consider the R&D cost. They are less bad at inference, but still much worse than even an A100.

I don't see how you can evaluate better and worse for training without doing so on cost basis. If it costs less and eventually finishes then it's better.

Re: How Meta trains large language models at scale

#39
post #28

Earlier quoted context omitted.

They do sell shovels, you can get Google TPUs on Google Cloud.

Exactly and they are still about 1/18ths as good at training llms as a H100. Maybe they are less than 1/18ths the cost, so google technically have a marginally better unit cost but i doubt it when you consider the R&D cost. They are less bad at inference, but still much worse than even an A100.

Also energy cost. 18 chips vs 1, it's probably costing a lot more to run 18

Re: How Meta trains large language models at scale

#40

Earlier quoted context omitted.

Wouldn't that be renting a shovel vs selling a shovel?

NVIDIA sells subscriptions...

I'm only aware of Nvidia AI Enterprise and that isn't required to run the GPU.

I think it's aimed at medium to large corporations.

Massive corporations such as Meta and OpenAI would build their own cloud and not rely on this.

The GPU really is a shovel, and can be used without any subscription.

Don't get me wrong, I want there to be competition with Nvidia, I want more access for open source and small players to run and train AI on competitive hardware at our own sites.

But no one is competing, no one has any idea what they're doing. Nvidia has no competition whatsoever, no one is even close.

This lets Nvidia get away with adding more vram onto an AI specific GPU and increase the price by 10x.

This lets Nvidia remove NVLink from current gen consumer cards like the 4090.

This lets Nvidia use their driver licence to prevent cloud platforms from offering consumers cards as a choice in datacenters.

If Nvidia had a shred of competition things would be much better.

Post reply on HN