Live data from Hacker News

Building Meta's GenAI infrastructure

engineering.fb.com

71–80 of 314 posts

Re: Building Meta's GenAI infrastructure

#71
post #20

Earlier quoted context omitted.

Isn't Google trying to do this with their TPUs?

I still, for the life of me, can't understand why Google doesn't just start selling their TPUs to everyone. Nvidia wouldn't be anywhere near their size if they only made H100s available through their DGX cloud, which is what Google is doing only making TPUs available through Google Cloud. Good hardware, good software support, and market is starving for performant competitors to the H100s (and soon B100s). Would sell…

It is an absolutely massive amount of work to turn something designed for your custom software stack and data centers (custom rack designs, water cooling, etc) into a COTS product that is plug-and-play; not just technically but also things like sales, support, etc. You are introducing a massive amount of new problems to solve and pay for. And the in-house designs like TPUs (or Meta's accelerators) are cost effective in part because they don't do that stuff at all. They would not be as cheap per unit of work if they had to also pay off all that other stuff. They also have had a very strong demand for TPUs internally which takes priority over GCP.

Re: Building Meta's GenAI infrastructure

#72

Honestly Meta is consistently one of the better companies at releasing tech stack info or just open sourcing, these kinds of articles are super fun

Do you find this informative?

Yes of course - it depends on what lens though. If you mean "I'm learning to build better from this" then no, but its very informative on Meta's own goals and mindset as well as real numbers that allow comparison to investment in other areas, etc. Also the point was mostly that Meta does publish a lot in the open - including actual open source tech stacks etc. They're reasonably good actors in this specific domain.

Re: Building Meta's GenAI infrastructure

#74
post #16

I wonder if Meta would ever try to compete with AWS / MSFT / GOOG for AI workloads

FB does not have the flywheel of running data centres - all three of those mentioned run hyper scale datacentres that they can then juice by “investing” billions in AI companies who then turn around and put those billions as revenue in the investors OpenAI takes money from MSFT and buys Azure services Anthropic takes Amazon money and buys AWS services (as do many robotics etc) I am fairly sure it’s not illegal but it…

Sounds like it's free equity at the very least

Re: Building Meta's GenAI infrastructure

#75

Earlier quoted context omitted.

I don't see how they're devaluing other people's AI products.

The angle is that by releasing cutting edge AI research to the public openly, the relative difference between open source models/tech and closed source tech shrinks. Whether or not you think the "value" of AI products is proportional to their performance gap vs the next closest thing or not is up to you. Very interesting PG essay I read recently talks about the opposite of this (Superlinear returns) where if you're h…

New Linux versions don't "salt the earth" for Windows.

Re: Building Meta's GenAI infrastructure

#77
post #35

How much are they paying for H100's? If they are paying $10k: 350,000 NVIDIA H100 x $10k = $3.5b

Significantly more than that; MFN pricing for NVIDIA DGX H100 (which has been getting priority supply allocation, so many have been suckered into buying them in order to get fast delivery) is ~$309k, while a basically equivalent HGX H100 system is ~$250k, coming to a price per GPU at the full server level being ~$31.5k. With Meta’s custom OCP systems integrating the SXM baseboards from NVIDIA, my guess is that their cost per GPU would be in the ~$23-$25k range.

Re: Building Meta's GenAI infrastructure

#78
post #66
post #55

Earlier quoted context omitted.

"My paycheck depends on this technology destroying every field producing cultural artifacts"

Said the butter churner, cotton ginner, and petrol pumper. I work in film. I've shot dozens of them the old fashioned way. I've always hated how labor, time, and cost intensive they are to make. Despite instructions from the luminaries to "just pick up a camera", the entire process is stone age. The field is extremely inequitable, full of nepotism and "who you know". Almost every starry-eyed film student winds up doi…

> And if you think any Jack or Jill can just come in and text prompt a whole movie, you're crazy. It's still hard work and a metric ton of good taste.

Yeah, I cant wait for ChuChuTV to get the best film Oscar /s.

Re: Building Meta's GenAI infrastructure

#80
post #20

Earlier quoted context omitted.

Isn't Google trying to do this with their TPUs?

I still, for the life of me, can't understand why Google doesn't just start selling their TPUs to everyone. Nvidia wouldn't be anywhere near their size if they only made H100s available through their DGX cloud, which is what Google is doing only making TPUs available through Google Cloud. Good hardware, good software support, and market is starving for performant competitors to the H100s (and soon B100s). Would sell…

The impression I got from this thread yesterday is that Google's having difficulty keeping up with the heavy internal demand for TPUs: https://news.ycombinator.com/item?id=39670121
Post reply on HN