350k H100 cards, around ten billion dollars just for the GPUs. Less if Nvidia gives a volume discount, which I imagine they do not.
It will be ironic if Meta sinks all this money into the new trend and finds out later that it has been a huge boondoggle, just as publishers followed Facebook's "guidance" on video being the future, subsequently gutting the talent pool and investing into video production and staff - only to find out it was all a total waste.
Building Meta's GenAI infrastructure
31–40 of 314 posts
Re: Building Meta's GenAI infrastructure
#32350k H100 cards, around ten billion dollars just for the GPUs. Less if Nvidia gives a volume discount, which I imagine they do not.
It will be ironic if Meta sinks all this money into the new trend and finds out later that it has been a huge boondoggle, just as publishers followed Facebook's "guidance" on video being the future, subsequently gutting the talent pool and investing into video production and staff - only to find out it was all a total waste.
Those GPUs are going to subsume the entire music, film, and gaming industries. And that's just to start.
Re: Building Meta's GenAI infrastructure
#33Interesting dig on IB. RoCE is the right solution since it is open standards and more importantly, available without a 52+ week lead time.
Re: Building Meta's GenAI infrastructure
#34350k H100 cards, around ten billion dollars just for the GPUs. Less if Nvidia gives a volume discount, which I imagine they do not.
It will be ironic if Meta sinks all this money into the new trend and finds out later that it has been a huge boondoggle, just as publishers followed Facebook's "guidance" on video being the future, subsequently gutting the talent pool and investing into video production and staff - only to find out it was all a total waste.
Re: Building Meta's GenAI infrastructure
#35Re: Building Meta's GenAI infrastructure
#36Earlier quoted context omitted.
What do you mean?
In pretty much every interview, Yann has talked about how important that AI infrastructure is open and distributed for the good of humanity, and how he wouldn't work for a company that wasn't open. Since Mark doesn't have an AI product to cannibalize, it's in his interest to devalue the AI products of others ("salting the earth").
Re: Building Meta's GenAI infrastructure
#37Total cluster they say will reach 350k H100, which at $30k street price is about $10b. In contrast, Microsoft is spending over $10b per quarter capex on cloud. That makes Zuck look conservative after his big loss on metaverse. https://www.datacenterdynamics.com/en/news/q3-2023-cloud-res...
What loss lol. Stop the fud
Re: Building Meta's GenAI infrastructure
#38I'd be great if they could invest in an alternative to nvidia -- then, in one fell swoop, destroy the moats of everyone in the industry.
A company moving away from Nvidia/CUDA while the field is developing so rapidly would result in that company falling behind. When (if) the rate of progress in the AI space slows, then perhaps the big players will have the breathing room to consider rethinking foundational components of their infrastructure. But even at that point, their massive investment in Nvidia will likely render this impractical. Nvidia decisive…
Re: Building Meta's GenAI infrastructure
#39float8 got a mention! x2 more FLOPs! Also xformers has 2:4 sparsity support now so another x2? Is Llama3 gonna use like float8 + 2:4 sparsity for the MLP, so 4x H100 float16 FLOPs? Pytorch has fp8 experimental support, whilst attention is still complex to do in float8 due to precision issues, so maybe attention is in float16, and RoPE / layernorms in float16 / float32, whilst everything else is float8?
Re: Building Meta's GenAI infrastructure
#40I'd be great if they could invest in an alternative to nvidia -- then, in one fell swoop, destroy the moats of everyone in the industry.
Isn't Google trying to do this with their TPUs?
Good hardware, good software support, and market is starving for performant competitors to the H100s (and soon B100s). Would sell like hotcakes.