Live data from Hacker News

Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)

cerebras.ai

21–30 of 100 posts

Re: Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)

#21
post #8

Maybe one day they’ll have an actual api that you can pay per token. Right now it’s the standard “talk to us” if you want to use it.

Huh? Just make an account, get your API key, and try out the free tier.. works for me. https://cloud.cerebras.ai

Yep, can confirm, I used their API just fine for Llama 4 Scout for weeks now.

Re: Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)

#22
post #19

I think it is too risky to build a company around the premise that someone won't soon solve the quadratic scaling issue. Especially, when that company involves creating ASICs. E.g.: https://arxiv.org/abs/2312.00752

Attention is not the primary inference bottleneck. For each token you have to load all of the weights (or activated weights) from memory. This is why Cerebras is fast: they have huge memory bandwidth.

Re: Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)

#23
post #18

are the Llama 4 issues fixed? what is it good at? coding is out of the window after the updated R1.

Yes, the issues were fixed ~1-2 weeks after release. It's a good "all-rounder" model, best compared to 4o. Good multilingual capabilities, even in languages not specifically highlighted. Fast to run inference on it. Code is not one of its strong suits at all.

Re: Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)

#24
post #3

> At over 2,500 t/s, Cerebras has set a world record for LLM inference speed on the 400B parameter Llama 4 Maverick model, the largest and most powerful in the Llama 4 family. This is incorrect. The unreleased Llama 4 Behemoth is the largest and most powerful in the Llama 4 family. As for the speed record, it seems important to keep it in context. That comparison is only for performance on 1 query, but it is well kno…

>Their hardware does inference with FP16, so they need ~20 of their CSE-3 chips to run this model. Care to explain? I don't see it.

CSE-3 chip has 44GB, which can hold 22B parameters in FP16.

400B parameters would need 18 chips. Then you need a bit more ram for other stuff

Re: Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)

#25
post #3

> At over 2,500 t/s, Cerebras has set a world record for LLM inference speed on the 400B parameter Llama 4 Maverick model, the largest and most powerful in the Llama 4 family. This is incorrect. The unreleased Llama 4 Behemoth is the largest and most powerful in the Llama 4 family. As for the speed record, it seems important to keep it in context. That comparison is only for performance on 1 query, but it is well kno…

> Also, Cerebras is the company that not only was saying that their hardware was not useful for inference until some time last year, but even partnered with Qualcomm with the claim that Qualcomm’s accelerators had a 10x price performance improvement over their things Mistral says they run Le Chat on Cerebras

Also perplexity

Re: Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)

#26
post #3

> At over 2,500 t/s, Cerebras has set a world record for LLM inference speed on the 400B parameter Llama 4 Maverick model, the largest and most powerful in the Llama 4 family. This is incorrect. The unreleased Llama 4 Behemoth is the largest and most powerful in the Llama 4 family. As for the speed record, it seems important to keep it in context. That comparison is only for performance on 1 query, but it is well kno…

> I have trouble getting excited about Cerebras. SRAM scaling is dead, so short of figuring out how to 3D stack their wafer scale chips AMD and TSMC are stacking SRAM on the chip scale. I imagine they could accomplish it at the wafer scale. It'll be neat if we can get hundreds of layers in time, like flash. Your analysis seems spot on to me.

Assume you meant Intel, rather than AMD?

Re: Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)

#27
post #4
post #2

Investors list include Altman and Ilya https://www.cerebras.ai/company

Their CEO is a felon who plead guilty to accounting fraud: https://milled.com/theinformation/cerebras-ceos-past-felony-... Experienced investors will not touch them: https://www.nbclosangeles.com/news/business/money-report/cer... I estimated last year that they can only produce about 300 chips per year and that is unlikely to change because there are far bigger customers for TSMC that are ahead of them in priority fo…

Openai wanted to buy them. G42 the largest player in middle east owne a big chunk. You are simply wrong about big investors not touching them but my guess is they will be bought soon by Meta or Apple.

Re: Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)

#28
post #26

Earlier quoted context omitted.

> I have trouble getting excited about Cerebras. SRAM scaling is dead, so short of figuring out how to 3D stack their wafer scale chips AMD and TSMC are stacking SRAM on the chip scale. I imagine they could accomplish it at the wafer scale. It'll be neat if we can get hundreds of layers in time, like flash. Your analysis seems spot on to me.

Assume you meant Intel, rather than AMD?

https://www.amd.com/en/products/processors/technologies/3d-v... and future developments.

Re: Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)

#29
post #3

> At over 2,500 t/s, Cerebras has set a world record for LLM inference speed on the 400B parameter Llama 4 Maverick model, the largest and most powerful in the Llama 4 family. This is incorrect. The unreleased Llama 4 Behemoth is the largest and most powerful in the Llama 4 family. As for the speed record, it seems important to keep it in context. That comparison is only for performance on 1 query, but it is well kno…

[deleted]

Re: Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)

#30
post #19

I think it is too risky to build a company around the premise that someone won't soon solve the quadratic scaling issue. Especially, when that company involves creating ASICs. E.g.: https://arxiv.org/abs/2312.00752

Yeah also strikes me as quite risky. Their gear seems very focused on llama family specifically.

Just takes one breakthrough and it's all different. See the recent diffusion style LLMs for example

Post reply on HN