Live data from Hacker News

Building Meta's GenAI infrastructure

engineering.fb.com

171–180 of 314 posts

Re: Building Meta's GenAI infrastructure

#171
post #111

This is great news for Nvidia and their stock, but are they sure the LLMs and image models will scale indefinitely? nature and biology has a preference for sigmoids. What if we find out that AGI requries different kinds of cpu capabilities

If anything, NVIDIA H100 GPUs are too general purpose! The optimal compute for AI training would be more specialised, but then would be efficient at only one NN architecture. Until we know what the best architecture is, the general purpose clusters remain a good strategy.

Re: Building Meta's GenAI infrastructure

#172

Earlier quoted context omitted.

Facebook has more datacenter space and power than Amazon, Google, and Microsoft -- possibly more than Amazon and Microsoft combined...

I have zero evidence, but this seems extremely unlikely. Do you have more than zero evidence?

Meta can use all their datacenter space while Amazon, Google, and Microsoft datacenter space is mostly rented.

Re: Building Meta's GenAI infrastructure

#173
post #59

Meta is still playing catch-up. Might be hard to believe but according to Reuters they've been trying to run AI workloads mostly on CPUs until 2022 and they had to pull the plug on the first iteration of their AI chip. https://www.reuters.com/technology/inside-metas-scramble-cat...

Definitely has some pr buzz and flex in the article. Now I see why.

Re: Building Meta's GenAI infrastructure

#174

float8 got a mention! x2 more FLOPs! Also xformers has 2:4 sparsity support now so another x2? Is Llama3 gonna use like float8 + 2:4 sparsity for the MLP, so 4x H100 float16 FLOPs? Pytorch has fp8 experimental support, whilst attention is still complex to do in float8 due to precision issues, so maybe attention is in float16, and RoPE / layernorms in float16 / float32, whilst everything else is float8?

Is it safe to assume this is the same float16 that exists in Apple m2 chips but not m1?

Re: Building Meta's GenAI infrastructure

#175

Earlier quoted context omitted.

Facebook has more datacenter space and power than Amazon, Google, and Microsoft -- possibly more than Amazon and Microsoft combined...

I don't think so, AWS hasn't disclosed this numbers, like datacenter spaces occupied, so how do you know.

I have mapped every AWS data center globally, and I worked at AWS.

Facebook publishes this data.

Re: Building Meta's GenAI infrastructure

#176

Earlier quoted context omitted.

People said the same thing when tensorflow was all the rage and pytorch was a side project. Granted, HW is much harder than SW, but I would not discount Meta's ability to displace NVIDIA entirely.

I don't think they could; nvidia has tons of talent, Meta would have to steal that. Meta doesn't do anything in either consumer or datacenter hardware that isn't for themselves either. Meta is a services company, their hardware is secondary and for their own usage.

meta has the Quest. It's not so bad that they're looking to create an LPU for their headset to offer local play.

Re: Building Meta's GenAI infrastructure

#177

Earlier quoted context omitted.

Facebook has more datacenter space and power than Amazon, Google, and Microsoft -- possibly more than Amazon and Microsoft combined...

Unless you've worked at Amazon, Microsoft, Google, and Facebook, or a whole bunch of datacenter providers, I'm not sure how you could make that claim. They don't really share that information freely, even in their stock reports. Heck I worked at Amazon and even then I couldn't tell you the total datacenter space, they don't even share it internally.

You can just map them all... I have. I also worked at AWS :)

Re: Building Meta's GenAI infrastructure

#178
post #143

Earlier quoted context omitted.

Yeah, and RoCE isn't single vendor. I'm not sure IB scales to the relevant cluster sizes, either.

Is NVLink just not scalable enough here?

I don't know. I haven't actually worked with IB in this specific space (or since before Nvidia acquired MLNX). My experience with RoCE/IB was for storage cluster backend in the late 2010s.

Re: Building Meta's GenAI infrastructure

#179

Earlier quoted context omitted.

How is it free equity? Spending money to invest it somewhere involves risks. You might recover some of it if the investment is valued by others, but there is no guarantee.

You do not need cash in hands to invest. Instead, you print your own money (AWS credit) and use that to drive up the valuation, because this money costs you nothing today. It might cost tomorrow though, when the company starts to use your services. However depending the deal structure they might not use all the credit, go belly up before credit is used or bought up by someone with real cash.

[deleted]

Re: Building Meta's GenAI infrastructure

#180

Earlier quoted context omitted.

FB does not have the flywheel of running data centres - all three of those mentioned run hyper scale datacentres that they can then juice by “investing” billions in AI companies who then turn around and put those billions as revenue in the investors OpenAI takes money from MSFT and buys Azure services Anthropic takes Amazon money and buys AWS services (as do many robotics etc) I am fairly sure it’s not illegal but it…

Such barter deals were also popular during the 00s Internet Bubble. Here more on the deals (2003): https://www.cnet.com/tech/services-and-software/aol-saga-ope... Popular names included AOL, Cisco, Yahoo, etc. Not sure if Amazon’s term sheets driving high valuation are nothing but AWS credits (Amazon’s own license to print money).

[deleted]
Post reply on HN