Live data from Hacker News

Meta MTIA v2 – Meta Training and Inference Accelerator

ai.meta.com

21–30 of 63 posts

Re: Meta MTIA v2 – Meta Training and Inference Accelerator

#21
post #13

I like the interactive 3D widget showing off the chip. Yep, that sure is a metal rectangle.

Really annoys me that the loading animation of these before-/after-images doesn't finish on firefox and that it won't let me drag the knob with the separator. ...no "Under the hood" for me.

Re: Meta MTIA v2 – Meta Training and Inference Accelerator

#22
post #9

Earlier quoted context omitted.

They are different architectures optimized for different things. From the Meta post: "This chip’s architecture is fundamentally focused on providing the right balance of compute, memory bandwidth, and memory capacity for serving ranking and recommendation models." Optimizing for ranking/recommendation models is very different from general purpose training/inference.

Yeah, it may fit their current workload perfectly, but it doesn't seem very future proof with the limited bandwidth. Given how fast ML is evolving these days I question if it makes sense to design and deploy a chip like this. I guess they do have a very large workload that will benefit immediately.

> Yeah, it may fit their current workload perfectly, but it doesn't seem very future proof

It’s custom silicon designed for a specific, known workload. It’s not designed to be a general purpose part or to be future proofed for unknown future applications.

When a new application comes along with new requirements, the teams will use their experience to create a new chip targeting that new application.

That’s the great part about custom silicon: You’re not hitting general specs for general applications that you may not even know about yet. You’re building one very specific thing to do a very specific job and do it very well.

Re: Meta MTIA v2 – Meta Training and Inference Accelerator

#24
post #9

Earlier quoted context omitted.

Yeah, it may fit their current workload perfectly, but it doesn't seem very future proof with the limited bandwidth. Given how fast ML is evolving these days I question if it makes sense to design and deploy a chip like this. I guess they do have a very large workload that will benefit immediately.

Don't mean to single you out at all, but I find this comment to be a great example of how the "ML Hype" is perceived by a certain segment folks in our industry. The development of this chip shows that it doesn't (and shouldn't!) matter to the ML teams at Meta how 'fast ML is evolving.' Indeed what it demonstrates is that a huge, global, trillion-dollar business has operationalized an existing ML technology to the ext…

To me, it’s bizarre to see the HPC mindset taking hold again after the cloud/commodity mindset dominated the last 16 years.

You don’t always need a Ferrari to go to the store

Re: Meta MTIA v2 – Meta Training and Inference Accelerator

#25
post #20

I find it weird that not everyone agree Meta and Facebook and social networks in general are doing some good the the society and our democracies; yet they manage to spend incredible amount of money/energy/time to develop solutions to problems we aren't exactly sure are worth solving…

If all this turns out to be useless, burning their cash for nothing seems like a great way to accelerate tech while going down. I guess that would actually be a positive thing.

Re: Meta MTIA v2 – Meta Training and Inference Accelerator

#26
post #24

Earlier quoted context omitted.

Don't mean to single you out at all, but I find this comment to be a great example of how the "ML Hype" is perceived by a certain segment folks in our industry. The development of this chip shows that it doesn't (and shouldn't!) matter to the ML teams at Meta how 'fast ML is evolving.' Indeed what it demonstrates is that a huge, global, trillion-dollar business has operationalized an existing ML technology to the ext…

To me, it’s bizarre to see the HPC mindset taking hold again after the cloud/commodity mindset dominated the last 16 years. You don’t always need a Ferrari to go to the store

WDYM by HPC mindset?

Re: Meta MTIA v2 – Meta Training and Inference Accelerator

#27
post #4

Pretty large increase in performance over v1, particularly in sparse workloads. Low power 25W Could use higher bandwidth memory if their workloads were more than recommendation engines.

First gen was 25W. The new one is 90W.

Ah, thanks for the correction.

Still relatively low compared to GPUs.

Re: Meta MTIA v2 – Meta Training and Inference Accelerator

#28
I thought MTIA v2 would use the mx formats https://arxiv.org/pdf/2302.08007.pdf, guess they were too far along in the process to get it in this time.

Still this looks like it would make for an amazing prosumer home ai setup. Could probably fit 12 accelerators on a wall outlet with change for a cpu, would have enough memory to serve a 2T model at 4bit and reasonable dense performance for small training runs and image stuff. Potentially not costing too much to make either without having to pay for cowos or hbm.

I'd definitely buy one if they ever decided to sell it and could keep the price under like $800/accelerator.

Re: Meta MTIA v2 – Meta Training and Inference Accelerator

#29
post #2

Intel Gaudi 3 has more interconnect bandwidth than this has memory bandwidth. By a lot. I guess they can't be fairly compared without knowing the TCO for each. I know in the past Google's TPU per-chip specs lagged Nvidia but the much lower TCO made them a slam dunk for Google's inference workloads. But this seems pretty far behind the state of the art. No FP8 either.

Only 48MB of SRAM on Gaudi 3 per die (96 MB across both) vs 256MB here maybe increases the memory bandwidth needs for Gaudi. Way different power consumption too.

Re: Meta MTIA v2 – Meta Training and Inference Accelerator

#30

I thought MTIA v2 would use the mx formats https://arxiv.org/pdf/2302.08007.pdf , guess they were too far along in the process to get it in this time. Still this looks like it would make for an amazing prosumer home ai setup. Could probably fit 12 accelerators on a wall outlet with change for a cpu, would have enough memory to serve a 2T model at 4bit and reasonable dense performance for small training runs and image…

I suppose it might, there are not a lot of details (what kind of sparsity for example?) about what they mean in terms of INT8 support - it could be MXINT8, or something else.

Glad someone was thinking the same thing I was though!

Post reply on HN