Live data from Hacker News

Building Meta's GenAI infrastructure

engineering.fb.com

151–160 of 314 posts

Re: Building Meta's GenAI infrastructure

#151

Earlier quoted context omitted.

A lot of the optimisation at this level is getting data into the right place at the right time, without killing the network. Its also a group effort to provide simple to use primitives that "normal" ML people can use, even if they've never used hyper scale clusters before. So you need a good scheduler, that understand dependencies (no, the k8s scheduler(s) are shit for this, plus it wont scale past 1k nodes without e…

Ok, but back to my main question, how do I get into this?

It looks more like an infra problem than ML. "Software architect"s mixed with devops/infra/sre people

Re: Building Meta's GenAI infrastructure

#152
post #88

Earlier quoted context omitted.

Proven technology, maybe, but proven product-market fit for the kinds of things Facebook is using it for? Their linked blog about AI features gives examples "AI stickers" and image editing... cool, but are these potential multi-billion dollar lifts to their existing business? I guess I'm skeptical it's worthwhile unless they're able to unseat ChatGPT with a market-leading general purpose assistant.

> unseat ChatGPT with a market-leading general purpose assistant. It's not impossible. The prediction from many(not that I believe it) is that over long run modelling tricks would become common knowledge and only thing that matters is compute and data, both of which Meta has. Also there could be a trend of LLMs for ads or feed recommendation in the future as they has large completely unstructured dataset per user acr…

Compute, data, and most importantly distribution/users.

IMO standalone AI companies like OpenAI might be successful by providing infrastructure to other companies, but I can’t imagine ChatGPT remaining #1 many years from now.

The web is still trending towards being a walled garden. Maybe not right now, but long term I think people will use whatever AI is most convenient which probably will be AI built into a giant company with established user base (FB, GOOG, MSFT, and Apple if they ever get around to launching - would love Siri 2.0 if it meant not needing to open the ChatGPT iOS app)

Re: Building Meta's GenAI infrastructure

#153
post #66
post #55

Earlier quoted context omitted.

"My paycheck depends on this technology destroying every field producing cultural artifacts"

Said the butter churner, cotton ginner, and petrol pumper. I work in film. I've shot dozens of them the old fashioned way. I've always hated how labor, time, and cost intensive they are to make. Despite instructions from the luminaries to "just pick up a camera", the entire process is stone age. The field is extremely inequitable, full of nepotism and "who you know". Almost every starry-eyed film student winds up doi…

> The field is extremely inequitable, full of nepotism and "who you know"

Maybe, but it's never been cheaper to make a movie.

I know someone with no connections and (almost) no money which in 4 years made multiple no. 1 box-office films (obviously not in US, in a smaller country) and then got picked up by Netflix.

Re: Building Meta's GenAI infrastructure

#154

Earlier quoted context omitted.

FB does not have the flywheel of running data centres - all three of those mentioned run hyper scale datacentres that they can then juice by “investing” billions in AI companies who then turn around and put those billions as revenue in the investors OpenAI takes money from MSFT and buys Azure services Anthropic takes Amazon money and buys AWS services (as do many robotics etc) I am fairly sure it’s not illegal but it…

Neither did AWS when they started. They were just building out data centers to run their little book website and decided to start selling the excess capacity. Meta could absolutely do the same, but in the short term, I think they find using that capacity more valuable than selling it.

> Neither did AWS when they started. They were just building out data centers to run their little book website and decided to start selling the excess capacity.

This is a myth. It simply isn't true. AWS was conceived as a greenfield business by its first CEO. Besides, S3 and SQS were the first AWS services; EC2 didn't appear till a few years later. And it wasn't built from excess Amazon server capacity; it was totally separate.

Re: Building Meta's GenAI infrastructure

#155
I think it’s always useful to pay attention to the history on stuff like this and it’s a rare pleasure to be able to give some pointers in the literature along with some color to those interested from first-hand experience.

I’d point the interested at the DLRM paper [1]: that was just after I left and I’m sad I missed it. FB got into disagg racks and SDN and stuff fairly early, and we already had half-U dual-socket SKUs with the SSD and (increasingly) even DRAM elsewhere in the rack in 2018, but we were doing huge NNs for recommenders and rankers even for then. I don’t know if this is considered proprietary so I’ll play it safe and just say that a click-prediction model on IG Stories in 2018 was on the order of a modest but real LLM today (at FP32!).

The crazy part is they were HOGWILD trained on Intel AVX-2, which is just wild to think about. When I was screwing around with CUDA kernels we were time sharing NVIDIA dev boxes, typically 2-4 people doing CUDA were splitting up a single card as late as maybe 2016. I was managing what was called “IGML Infra” when I left and was on a first-name basis with the next-gen hardware people and any NVIDIA deal was still so closely guarded I didn’t hear more than rumors about GPUs for training let alone inference.

350k Hopper this year, Jesus. Say what you want about Meta but don’t say they can’t pour concrete and design SKUs on a dime: best damned infrastructure folks in the game pound-for-pound to this day.

The talk by Thomas “tnb” Bredillet in particular I’d recommend: one of the finest hackers, mathematicians, and humans I’ve ever had the pleasure to know.

[1] https://arxiv.org/pdf/1906.00091.pdf

[2] https://arxiv.org/pdf/2108.09373.pdf

[3] https://engineering.fb.com/2022/10/18/open-source/ocp-summit...

[4] https://youtu.be/lQlIwWVlPGo?si=rRbRUAXX7aM0UcVO

Re: Building Meta's GenAI infrastructure

#156
post #101

float8 got a mention! x2 more FLOPs! Also xformers has 2:4 sparsity support now so another x2? Is Llama3 gonna use like float8 + 2:4 sparsity for the MLP, so 4x H100 float16 FLOPs? Pytorch has fp8 experimental support, whilst attention is still complex to do in float8 due to precision issues, so maybe attention is in float16, and RoPE / layernorms in float16 / float32, whilst everything else is float8?

Is there float8 support in any common CPU intrinsics? It sounds interesting but curious what will be the impact if any on CPU inference.

[deleted]

Re: Building Meta's GenAI infrastructure

#158

Earlier quoted context omitted.

Ok, but back to my main question, how do I get into this?

It looks more like an infra problem than ML. "Software architect"s mixed with devops/infra/sre people

Well since I'm not a ML engineer of any kind - that's good!

Re: Building Meta's GenAI infrastructure

#159

Earlier quoted context omitted.

The big difference is that back then, anyone with a consumer-level computer in their bedroom could turn it into a server and be a first-class citizen on the Internet. With generative AI, models will be controlled by a handful of giant corporations who have the enormous corpuses (of dubious provenance) and compute ability to train them. So it will be like last time, but even worse.

You can run ComfyUI and AnimateDiff on your PC. If you haven't checked them out, please do. And there are other angles to consider. Apple, for one, is expressly interested in not becoming a thin client to cloud AI. They're baking a lot of inference power into their chips. If the creative class don't need their devices, that doesn't bode well for them...

Running local models isn't the same as being able to train them from scratch yourself on a corpus of your own choosing.

Re: Building Meta's GenAI infrastructure

#160

Earlier quoted context omitted.

You can run ComfyUI and AnimateDiff on your PC. If you haven't checked them out, please do. And there are other angles to consider. Apple, for one, is expressly interested in not becoming a thin client to cloud AI. They're baking a lot of inference power into their chips. If the creative class don't need their devices, that doesn't bode well for them...

Running local models isn't the same as being able to train them from scratch yourself on a corpus of your own choosing.

There are so many ways to do exactly this too!

FakeYou, CivitAi, WeightsGg, Comflowy, ... -- there are tons of vibrant communities to teach you everything you need to know. The tools are open source, free to use, and accessible.

This isn't hard at all once you dive in.

Post reply on HN