Live data from Hacker News

Building Meta's GenAI infrastructure

engineering.fb.com

11–20 of 314 posts

Re: Building Meta's GenAI infrastructure

#11
post #8

I know we won't get it this from FB, but I'd be really interested to see how the relationship of compute power to engineering hours scales. They mention custom building as much as they can. If FB magically has the option to 10x the compute power, would they need to re-engineer the whole stack? What about 100x? Is each of these re-writes just a re-write, or is it a whole order of magnitude more complex? My technical u…

I'm not 100% sure but I would.make an educated guess that that cluster in the first image for example is a sample of scalable clusters, so throwing more hardware at it could bring improvements but sooner or later the cost to improvements will call for an optimization or rewrite as you call it, so a bit of both usually. It seems a bit of a balancing act really!

Re: Building Meta's GenAI infrastructure

#12

I'd be great if they could invest in an alternative to nvidia -- then, in one fell swoop, destroy the moats of everyone in the industry.

A company moving away from Nvidia/CUDA while the field is developing so rapidly would result in that company falling behind. When (if) the rate of progress in the AI space slows, then perhaps the big players will have the breathing room to consider rethinking foundational components of their infrastructure. But even at that point, their massive investment in Nvidia will likely render this impractical. Nvidia decisively won the AI hardware lottery, and that's why it's worth trillions.

Re: Building Meta's GenAI infrastructure

#13
Total cluster they say will reach 350k H100, which at $30k street price is about $10b.

In contrast, Microsoft is spending over $10b per quarter capex on cloud.

That makes Zuck look conservative after his big loss on metaverse.

https://www.datacenterdynamics.com/en/news/q3-2023-cloud-res...

Re: Building Meta's GenAI infrastructure

#14

Total cluster they say will reach 350k H100, which at $30k street price is about $10b. In contrast, Microsoft is spending over $10b per quarter capex on cloud. That makes Zuck look conservative after his big loss on metaverse. https://www.datacenterdynamics.com/en/news/q3-2023-cloud-res...

What loss lol. Stop the fud

Re: Building Meta's GenAI infrastructure

#17

Total cluster they say will reach 350k H100, which at $30k street price is about $10b. In contrast, Microsoft is spending over $10b per quarter capex on cloud. That makes Zuck look conservative after his big loss on metaverse. https://www.datacenterdynamics.com/en/news/q3-2023-cloud-res...

That's a weird comparison. The GPU is only a part of the capex: there's the rest of the servers and racks, the networking, as well as the buildings/cooling systems to support that.

Re: Building Meta's GenAI infrastructure

#18

Yann wants to be open and Mark seems happy to salt the earth.

What do you mean?

In pretty much every interview, Yann has talked about how important that AI infrastructure is open and distributed for the good of humanity, and how he wouldn't work for a company that wasn't open. Since Mark doesn't have an AI product to cannibalize, it's in his interest to devalue the AI products of others ("salting the earth").

Re: Building Meta's GenAI infrastructure

#19

I'd be great if they could invest in an alternative to nvidia -- then, in one fell swoop, destroy the moats of everyone in the industry.

Except that "one fell swoop" would realistically be 20+ years of research and development from the top minds in the semiconductor industry.
Post reply on HN