My mind still boggles that a BBS+ads company would think it needs to design its own chips.
Meta MTIA v2 – Meta Training and Inference Accelerator
11–20 of 63 posts
Re: Meta MTIA v2 – Meta Training and Inference Accelerator
#12Earlier quoted context omitted.
They are different architectures optimized for different things. From the Meta post: "This chip’s architecture is fundamentally focused on providing the right balance of compute, memory bandwidth, and memory capacity for serving ranking and recommendation models." Optimizing for ranking/recommendation models is very different from general purpose training/inference.
Yeah, it may fit their current workload perfectly, but it doesn't seem very future proof with the limited bandwidth. Given how fast ML is evolving these days I question if it makes sense to design and deploy a chip like this. I guess they do have a very large workload that will benefit immediately.
The development of this chip shows that it doesn't (and shouldn't!) matter to the ML teams at Meta how 'fast ML is evolving.'
Indeed what it demonstrates is that a huge, global, trillion-dollar business has operationalized an existing ML technology to the extent that they can invest into, and deploy, customized hardware for solving a business problem.
How ML "evolves" is irrelevant. They have a system which solves their problem, and they're investing in it.
Re: Meta MTIA v2 – Meta Training and Inference Accelerator
#13Re: Meta MTIA v2 – Meta Training and Inference Accelerator
#14My mind still boggles that a BBS+ads company would think it needs to design its own chips.
Re: Meta MTIA v2 – Meta Training and Inference Accelerator
#15I can only imagine the lack of fear Jensen experiences when reading this.
Re: Meta MTIA v2 – Meta Training and Inference Accelerator
#16My mind still boggles that a BBS+ads company would think it needs to design its own chips.
Re: Meta MTIA v2 – Meta Training and Inference Accelerator
#17Intel Gaudi 3 has more interconnect bandwidth than this has memory bandwidth. By a lot. I guess they can't be fairly compared without knowing the TCO for each. I know in the past Google's TPU per-chip specs lagged Nvidia but the much lower TCO made them a slam dunk for Google's inference workloads. But this seems pretty far behind the state of the art. No FP8 either.
LPDDR5 vs HBMe2. I'm guessing there's a 2-5x price difference between those, but even so it's an interesting choice, I don't know any other accelerators which spec DDR. But yeah, without exact TCO numbers it's hard to compare exactly.
Re: Meta MTIA v2 – Meta Training and Inference Accelerator
#18Earlier quoted context omitted.
Yeah, it may fit their current workload perfectly, but it doesn't seem very future proof with the limited bandwidth. Given how fast ML is evolving these days I question if it makes sense to design and deploy a chip like this. I guess they do have a very large workload that will benefit immediately.
Don't mean to single you out at all, but I find this comment to be a great example of how the "ML Hype" is perceived by a certain segment folks in our industry. The development of this chip shows that it doesn't (and shouldn't!) matter to the ML teams at Meta how 'fast ML is evolving.' Indeed what it demonstrates is that a huge, global, trillion-dollar business has operationalized an existing ML technology to the ext…
You've gotta learn to walk before you can run
Re: Meta MTIA v2 – Meta Training and Inference Accelerator
#19Still seems pretty primitive. Very cool though. I can only imagine the lack of fear Jensen experiences when reading this.