Earlier quoted context omitted.
Wouldn't that be renting a shovel vs selling a shovel?
NVIDIA sells subscriptions...
How Meta trains large language models at scale
151–160 of 213 posts
Re: How Meta trains large language models at scale
#152OK this was a bit funny: Top HW failure modes: * GPU falling off the bus I honestly thought "do they mean GPUs falling off a bus entering the data center" and then realized its actually the connectivity, as they mention in the next line GPUs falling off: In this case, GPUs are not detected by the host on PCIe.
Re: How Meta trains large language models at scale
#153Earlier quoted context omitted.
> How do they sanitize PII? I can't comment on how things like faces get used, but in my experience, PII at Meta is inaccessible by default. Unless you're impersonating a user on the platform (to access what PII they can see), you have to request special access for logs or database columns that contain so much as user IDs, otherwise the data simply won't show up when you query for it. This is baked into the infrastru…
For a convenient definition of PII. Isn’t everything a user does in aggregate PII?
Things that aren't PII aren't "convenient" definitions. Doesn't mean everything that isn't PII is fine to share. It's like saying a kidnapping isn't a murder. That's not a convenient definition of murder; it's just a different thing. We shouldn't start talking like witch hunters as soon as we encounter a situation that we haven't memorised a reasonable response to. We should be able to respond reasonably to new situations.
Re: How Meta trains large language models at scale
#154OK this was a bit funny: Top HW failure modes: * GPU falling off the bus I honestly thought "do they mean GPUs falling off a bus entering the data center" and then realized its actually the connectivity, as they mention in the next line GPUs falling off: In this case, GPUs are not detected by the host on PCIe.
I'm wondering if we could prompt llama3 with the above statement. What kind of response would it give?
Re: How Meta trains large language models at scale
#155Earlier quoted context omitted.
> How do they sanitize PII? I can't comment on how things like faces get used, but in my experience, PII at Meta is inaccessible by default. Unless you're impersonating a user on the platform (to access what PII they can see), you have to request special access for logs or database columns that contain so much as user IDs, otherwise the data simply won't show up when you query for it. This is baked into the infrastru…
For a convenient definition of PII. Isn’t everything a user does in aggregate PII?
Obvious examples: data that easily identifies a person (Photo, name, number, UUID, etc)
Thats trivial to block. Where it gets harder is stuff that on it's own isn't PII, but combined with another source, would be
For example, aggregating public comments on a celeb's post. (ie stripping out usernames and likes and assigning a new UUID to each person.) For a single post, thats good enough. You're very unlikley to be able to identify a single person.
But over multiple posts, thats where it gets tricky.
As with large companies, the process for getting permission to use that kind of data is righty difficult, so it often doesn't get used like that.
Re: How Meta trains large language models at scale
#156Posts like this underscore why the smart money is betting on Google as the long term AI winner. Meta, Microsoft, OpenAI, etc. are trying to address problems with consumer video cards and spending billions to try and out bid each other to win Nvidia's favor - while Google is on their 6th generation of custom silicon. Literally the only thing that can stop Google now is the fact they keep bringing Microsoft and Oracle…
Yes, but you are buying access to tested, supported units that are proven to work, don't require custom software, and are almost plug an go. When its time to upgrade, its not that costly.
Designing, fabricating and deploying your own silicon is Expensive, creating software support for it, also more expense. THen there is the opportunity cost of having to optimise the software stack your self.
You're exchanging a large capex, for a similar sized capex plus a fuckton of opex as well.
Re: How Meta trains large language models at scale
#157Earlier quoted context omitted.
Can't agree. This is like saying $popularApp will fail because they buy expensive hosting at AWS. Rubbish they will fail because the product didn't fit the market, if they're successful they'll have money to buy servers and colo then drive down cost. If they succeed it will be in large part due to the fact they spent thier capital and more importantly time on code/engineers rather than servers. Right now companies ar…
> This is like saying $popularApp will fail because they buy expensive hosting at AWS. For any given mobile app startup, AWS is effectively infinite. The more money you throw at it the more doodads you get back. Nvidia's supply chain is not infinite and is the bottle neck for all the non-Google players to fight over.
Re: How Meta trains large language models at scale
#158Would be nice to read how do they collect/prepare data for training. Which data sources? How much of Meta users data (fb, instagram… etc). How do they sanitize PII?
In the paper covering the original Llama they explicitly list their data sources in table 1 - including saying that they pretrained on the somewhat controversial books3 dataset.
The paper for Llama 2 also explicitly says they don't take data from Meta's products and services; and that they filter out data from sites known to contain a lot of PII. Although it is more coy about precisely what data sources they used, like many such papers are.
Re: How Meta trains large language models at scale
#159OK this was a bit funny: Top HW failure modes: * GPU falling off the bus I honestly thought "do they mean GPUs falling off a bus entering the data center" and then realized its actually the connectivity, as they mention in the next line GPUs falling off: In this case, GPUs are not detected by the host on PCIe.
> GPU falling off the bus I'm wondering if we could prompt llama3 with the above statement. What kind of response would it give?
The infamous "GPU falling off the bus" issue!
This problem typically occurs when a graphics processing unit (GPU) is not properly seated or connected to its expansion slot, such as PCIe, on a motherboard.
Here are some troubleshooting steps to help resolve the issue:
(numbered list of steps or options follows)
Tested on Llama 3 Instruct 7B Q8_0, because that one fits entirely on my GPU.
Re: How Meta trains large language models at scale
#160OK this was a bit funny: Top HW failure modes: * GPU falling off the bus I honestly thought "do they mean GPUs falling off a bus entering the data center" and then realized its actually the connectivity, as they mention in the next line GPUs falling off: In this case, GPUs are not detected by the host on PCIe.
Actually, they "fell" off the truck: https://www.theverge.com/2021/11/6/22767046/someone-stole-sh...