Live data from Hacker News

How Meta trains large language models at scale

engineering.fb.com

71–80 of 213 posts

Re: How Meta trains large language models at scale

#71
> Since we did not have time to change the cooling infrastructure, we had to remain in an air-cooled environment. The mechanical and thermal designs had to change to accommodate this, and that triggered a validation cycle to support a large-scale deployment.

> All of these hardware-related changes were challenging because we had to find a solution that fit within the existing resource constraints, with a very small degree of freedom to change and meet a tight schedule.

Seems like the time constraints put into the team impacted the overall quality of the model.

Re: How Meta trains large language models at scale

#72
post #42

Earlier quoted context omitted.

Ever since Apple did it everyone has leaped on board. Let's see how things pan out for everyone...

Google introduced their first TPU in 2015...? Long before Apple taped out their first silicon.

if we're talking custom silicon, Google acquired Motorola in 2011, and Apple acquired PA semi in 2008.

The idea is obvious to everybody in the industry, it's a question of money and motivation.

Re: How Meta trains large language models at scale

#73
post #4

Posts like this underscore why the smart money is betting on Google as the long term AI winner. Meta, Microsoft, OpenAI, etc. are trying to address problems with consumer video cards and spending billions to try and out bid each other to win Nvidia's favor - while Google is on their 6th generation of custom silicon. Literally the only thing that can stop Google now is the fact they keep bringing Microsoft and Oracle…

smart money has a diversified portfolio and isn't betting on any one winner and has invested in all of them, and then some.

Re: How Meta trains large language models at scale

#74

Earlier quoted context omitted.

I think it's likely Nvidia's GPU's, many of which are $50,000+ for a single unit, far surpass Google's custom silicon otherwise why wouldn't Google be selling shovels like Nvidia? If Google had a better chip, or even a chip that was close, they would sell it to anyone and everyone. From a quick search I can see Google's custom chips are 15x to 30x slower to train AI compared to Nvidia's current latest gen AI specific…

We have almost 400 H100's sitting idle. I wonder how many other companies are buying millions of dollars worth of these chips with the hopes of them being used, but aren't being utilized?

the world would love to buy time on your idle H100's if you're selling.

Re: How Meta trains large language models at scale

#75
post #26
post #4

Posts like this underscore why the smart money is betting on Google as the long term AI winner. Meta, Microsoft, OpenAI, etc. are trying to address problems with consumer video cards and spending billions to try and out bid each other to win Nvidia's favor - while Google is on their 6th generation of custom silicon. Literally the only thing that can stop Google now is the fact they keep bringing Microsoft and Oracle…

Can't agree. This is like saying $popularApp will fail because they buy expensive hosting at AWS. Rubbish they will fail because the product didn't fit the market, if they're successful they'll have money to buy servers and colo then drive down cost. If they succeed it will be in large part due to the fact they spent thier capital and more importantly time on code/engineers rather than servers. Right now companies ar…

[deleted]

Re: How Meta trains large language models at scale

#76
post #38
post #28

Earlier quoted context omitted.

Exactly and they are still about 1/18ths as good at training llms as a H100. Maybe they are less than 1/18ths the cost, so google technically have a marginally better unit cost but i doubt it when you consider the R&D cost. They are less bad at inference, but still much worse than even an A100.

I don't see how you can evaluate better and worse for training without doing so on cost basis. If it costs less and eventually finishes then it's better.

This assumes that you can linearly scale up the number of TPUs to get equal performance to Nvidia cards for less cost. Like most things distributed, this is unlikely to be the case.

Re: How Meta trains large language models at scale

#77
post #45

How will Meta leverage LLMs at scale to drive revenue? It's not clear.

I still believe that a VR future is coming once the technology commoditises i.e. costs come down 10x and we have 3090 level GPUs in the headset. At that point we will have photo-realistic experiences like concerts etc that anyone can afford.

And at that point having a lot of LLM based avatars that can help "fill in the space" will be valuable.

Re: How Meta trains large language models at scale

#79

OK this was a bit funny: Top HW failure modes: * GPU falling off the bus I honestly thought "do they mean GPUs falling off a bus entering the data center" and then realized its actually the connectivity, as they mention in the next line GPUs falling off: In this case, GPUs are not detected by the host on PCIe.

Brings a whole new meaning to bus factor

Re: How Meta trains large language models at scale

#80
post #38
post #28

Earlier quoted context omitted.

Exactly and they are still about 1/18ths as good at training llms as a H100. Maybe they are less than 1/18ths the cost, so google technically have a marginally better unit cost but i doubt it when you consider the R&D cost. They are less bad at inference, but still much worse than even an A100.

I don't see how you can evaluate better and worse for training without doing so on cost basis. If it costs less and eventually finishes then it's better.

Time is money. You might be a lab with long queues to train, leaving expensive staff twiddling their thumbs.
Post reply on HN