Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)
1–10 of 100 posts
Re: Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)
#2Re: Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)
#3This is incorrect. The unreleased Llama 4 Behemoth is the largest and most powerful in the Llama 4 family.
As for the speed record, it seems important to keep it in context. That comparison is only for performance on 1 query, but it is well known that people run potentially hundreds of queries in parallel to get their money out of the hardware. If you aggregate the tokens per second across all simultaneous queries to get the total throughput for comparison, I wonder if it will still look so competitive in absolute performance.
Also, Cerebras is the company that not only was saying that their hardware was not useful for inference until some time last year, but even partnered with Qualcomm with the claim that Qualcomm’s accelerators had a 10x price performance improvement over their things:
https://www.cerebras.ai/press-release/cerebras-qualcomm-anno...
Their hardware does inference with FP16, so they need ~20 of their CSE-3 chips to run this model. Each one costs ~$2 million, so that is $40 million. The DGX B200 that they used for their comparison costs ~$500,000:
https://wccftech.com/nvidia-blackwell-dgx-b200-price-half-a-...
You only need 1 DGX B200 to run Llama 4 Maverick. You could buy ~80 of them for the price it costs to buy enough Cerebras hardware to run Llama 4 Maverick.
Their latencies are impressive, but beyond a certain point, throughput is what counts and they don’t really talk about their throughput numbers. I suspect the cost to performance ratio is terrible for throughput numbers. It certainly is terrible for latency numbers. That is what they are not telling people.
Finally, I have trouble getting excited about Cerebras. SRAM scaling is dead, so short of figuring out how to 3D stack their wafer scale chips, during fabrication at TSMC, or designing round chips, they have a dead end product since it relies on using an entire wafer to be able to throw SRAM at problems. Nvidia, using DRAM, is far less reliant on SRAM and can use more silicon for compute, which is still shrinking.
Re: Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)
#4Investors list include Altman and Ilya https://www.cerebras.ai/company
https://milled.com/theinformation/cerebras-ceos-past-felony-...
Experienced investors will not touch them:
https://www.nbclosangeles.com/news/business/money-report/cer...
I estimated last year that they can only produce about 300 chips per year and that is unlikely to change because there are far bigger customers for TSMC that are ahead of them in priority for capacity. Their technology is interesting, but it is heavily reliant on SRAM and SRAM scaling is dead. Unless they get a foundry to stack layers for their wafer scale chips or design a round chip, they are unlikely to be able to improve their technology very much past the CSE-3. Compute might somewhat increase in the CSE-4 if there is one, but memory will not increase much if at all.
I doubt the investors will see a return on investment.
Re: Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)
#5Re: Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)
#6Investors list include Altman and Ilya https://www.cerebras.ai/company
Re: Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)
#7Investors list include Altman and Ilya https://www.cerebras.ai/company
Their CEO is a felon who plead guilty to accounting fraud: https://milled.com/theinformation/cerebras-ceos-past-felony-... Experienced investors will not touch them: https://www.nbclosangeles.com/news/business/money-report/cer... I estimated last year that they can only produce about 300 chips per year and that is unlikely to change because there are far bigger customers for TSMC that are ahead of them in priority fo…
Re: Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)
#8Re: Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)
#9> At over 2,500 t/s, Cerebras has set a world record for LLM inference speed on the 400B parameter Llama 4 Maverick model, the largest and most powerful in the Llama 4 family. This is incorrect. The unreleased Llama 4 Behemoth is the largest and most powerful in the Llama 4 family. As for the speed record, it seems important to keep it in context. That comparison is only for performance on 1 query, but it is well kno…
Emphasis mine.
Behemoth may become the largest and most powerful llama model, but right now it's nothing but vaporware. Maverick is currently the largest and more powerful llama model today (and if I had to bet, my money would be on Meta discarding Llama4 Behemoth entirely it eventually without having released it, and moving on to the next version number).
Re: Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)
#10> At over 2,500 t/s, Cerebras has set a world record for LLM inference speed on the 400B parameter Llama 4 Maverick model, the largest and most powerful in the Llama 4 family. This is incorrect. The unreleased Llama 4 Behemoth is the largest and most powerful in the Llama 4 family. As for the speed record, it seems important to keep it in context. That comparison is only for performance on 1 query, but it is well kno…
AMD and TSMC are stacking SRAM on the chip scale. I imagine they could accomplish it at the wafer scale. It'll be neat if we can get hundreds of layers in time, like flash.
Your analysis seems spot on to me.