Live data from Hacker News

How Meta trains large language models at scale

engineering.fb.com

121–130 of 213 posts

Re: How Meta trains large language models at scale

#121
post #4

Posts like this underscore why the smart money is betting on Google as the long term AI winner. Meta, Microsoft, OpenAI, etc. are trying to address problems with consumer video cards and spending billions to try and out bid each other to win Nvidia's favor - while Google is on their 6th generation of custom silicon. Literally the only thing that can stop Google now is the fact they keep bringing Microsoft and Oracle…

this was the same argument that was presented a decade ago on why Google was supposed to win the cloud because their internal infra was miles ahead of Amazon and Microsoft.

Yet here we are. Will the consumer video cards get cheaper and better faster or will Google's directors' infighting stop first?

Re: How Meta trains large language models at scale

#122

OK this was a bit funny: Top HW failure modes: * GPU falling off the bus I honestly thought "do they mean GPUs falling off a bus entering the data center" and then realized its actually the connectivity, as they mention in the next line GPUs falling off: In this case, GPUs are not detected by the host on PCIe.

A GPU falling off the bus would be one mega flop

Re: How Meta trains large language models at scale

#123

OK this was a bit funny: Top HW failure modes: * GPU falling off the bus I honestly thought "do they mean GPUs falling off a bus entering the data center" and then realized its actually the connectivity, as they mention in the next line GPUs falling off: In this case, GPUs are not detected by the host on PCIe.

Actually, they "fell" off the truck:

https://www.theverge.com/2021/11/6/22767046/someone-stole-sh...

Re: How Meta trains large language models at scale

#124
post #84

Would be nice to read how do they collect/prepare data for training. Which data sources? How much of Meta users data (fb, instagram… etc). How do they sanitize PII?

> How do they sanitize PII?

I can't comment on how things like faces get used, but in my experience, PII at Meta is inaccessible by default.

Unless you're impersonating a user on the platform (to access what PII they can see), you have to request special access for logs or database columns that contain so much as user IDs, otherwise the data simply won't show up when you query for it. This is baked into the infrastructure layer, so I doubt the GenAI teams are using something else.

Re: How Meta trains large language models at scale

#125

OK this was a bit funny: Top HW failure modes: * GPU falling off the bus I honestly thought "do they mean GPUs falling off a bus entering the data center" and then realized its actually the connectivity, as they mention in the next line GPUs falling off: In this case, GPUs are not detected by the host on PCIe.

A GPU falling off the bus would be one mega flop

The audience in the back goes clap clap clap, chapeau bas.

Re: How Meta trains large language models at scale

#126

Yikes, the little Infiniband+A100 cluster I installed for my previous company seemed useful at the time (12 GPUs) and that was at a cost of around $300k. With LLMs it feels like game over for non-cloud applications if you are not a mega-corp.

Well, yes, but Not all models need to be "super large." Smaller models, specialized in specific tasks, working together - and then reporting to a slightly larger model is the way to go.

Think of everything being connected to a "Home Computer" in those "Future House of 2020" videos that were out there in 70s or what not.

Another example (very rough) would be something like "Weather data gets to a small model via an API, model looks at it, updates the home dashboard, also sees if there's any alerts, if so, adds x or y to home dashboard appropriately as to what it thinks best."

We can probably achieve the latter example today. (without any significant 'coding' on anyone's part except the API owner)

Re: How Meta trains large language models at scale

#127
post #45

How will Meta leverage LLMs at scale to drive revenue? It's not clear.

LLMs aren't a monetizable product themselves. For the foreseeable future, that will always be ads. LLMs (and VR) are just big bets on getting ahead of future technology.

Re: How Meta trains large language models at scale

#129
post #118

Earlier quoted context omitted.

Don’t forget Apple’s Private Compute Cloud - built on top of Apple Silicon.

Models trained on Google TPUs according to Reuters [0]. Does anyone know the "technical document" the news article references? [0] https://www.reuters.com/technology/artificial-intelligence/h...

https://machinelearning.apple.com/research/introducing-apple...

Re: How Meta trains large language models at scale

#130

Earlier quoted context omitted.

Are you suggesting Meta or Google - who stand to save billions - won't be able to get top performance from their custom chips because their tooling/hardware won't support CUDA?

No. I’m suggesting they won’t because IP like the mellanox treasure chest they acquired is ridiculously difficult to develop and Nvidia has aggressively exploited it, along with their other already advanced IP in the space of their -core business-. I understand, especially amongst googlers, there’s a belief there are no others smarter than a googler. But it’s simply not the case. Nvidia is excellent at their core com…

The bar for success for Google and Meta is much lower than Nvidia - at least for internal usage. Any dollar amount that Google saves on CapEx or OpEx by using custom silicon instead of buying Nvidia helps bring down the cost of revenue. They don't have to match Nvidia on raw performance, and can aim at being better at performance per watt or performance per dollar (TCO) for larger workloads, and IIRC, Google is already doing for some internal inferencing tasks.

> I would also note that at this phase of a cycle in tech trying to save billions takes your eye off the prize

Big Tech companies are conglomerate-ish and can multitask. The search engine folk aren't pushing stuff back onto the backlog to put out fires delaying chip tape-out, and I bet the respective CEOs aren't burning braincycles micromanaging silicon development either; directors 2-3 rungs below the C-suite can motivate and execute on such an undertaking. The answer to "I need a budget of $300M in order to save the company $5-15B over 3 years" is "How soon can you start?"

Post reply on HN