Maybe one day they’ll have an actual api that you can pay per token. Right now it’s the standard “talk to us” if you want to use it.
Huh? Just make an account, get your API key, and try out the free tier.. works for me. https://cloud.cerebras.ai
Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)
21–30 of 100 posts
Re: Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)
#22I think it is too risky to build a company around the premise that someone won't soon solve the quadratic scaling issue. Especially, when that company involves creating ASICs. E.g.: https://arxiv.org/abs/2312.00752
Re: Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)
#23are the Llama 4 issues fixed? what is it good at? coding is out of the window after the updated R1.
Re: Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)
#24> At over 2,500 t/s, Cerebras has set a world record for LLM inference speed on the 400B parameter Llama 4 Maverick model, the largest and most powerful in the Llama 4 family. This is incorrect. The unreleased Llama 4 Behemoth is the largest and most powerful in the Llama 4 family. As for the speed record, it seems important to keep it in context. That comparison is only for performance on 1 query, but it is well kno…
>Their hardware does inference with FP16, so they need ~20 of their CSE-3 chips to run this model. Care to explain? I don't see it.
400B parameters would need 18 chips. Then you need a bit more ram for other stuff
Re: Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)
#25> At over 2,500 t/s, Cerebras has set a world record for LLM inference speed on the 400B parameter Llama 4 Maverick model, the largest and most powerful in the Llama 4 family. This is incorrect. The unreleased Llama 4 Behemoth is the largest and most powerful in the Llama 4 family. As for the speed record, it seems important to keep it in context. That comparison is only for performance on 1 query, but it is well kno…
> Also, Cerebras is the company that not only was saying that their hardware was not useful for inference until some time last year, but even partnered with Qualcomm with the claim that Qualcomm’s accelerators had a 10x price performance improvement over their things Mistral says they run Le Chat on Cerebras
Re: Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)
#26> At over 2,500 t/s, Cerebras has set a world record for LLM inference speed on the 400B parameter Llama 4 Maverick model, the largest and most powerful in the Llama 4 family. This is incorrect. The unreleased Llama 4 Behemoth is the largest and most powerful in the Llama 4 family. As for the speed record, it seems important to keep it in context. That comparison is only for performance on 1 query, but it is well kno…
> I have trouble getting excited about Cerebras. SRAM scaling is dead, so short of figuring out how to 3D stack their wafer scale chips AMD and TSMC are stacking SRAM on the chip scale. I imagine they could accomplish it at the wafer scale. It'll be neat if we can get hundreds of layers in time, like flash. Your analysis seems spot on to me.
Re: Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)
#27Investors list include Altman and Ilya https://www.cerebras.ai/company
Their CEO is a felon who plead guilty to accounting fraud: https://milled.com/theinformation/cerebras-ceos-past-felony-... Experienced investors will not touch them: https://www.nbclosangeles.com/news/business/money-report/cer... I estimated last year that they can only produce about 300 chips per year and that is unlikely to change because there are far bigger customers for TSMC that are ahead of them in priority fo…
Re: Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)
#28Earlier quoted context omitted.
> I have trouble getting excited about Cerebras. SRAM scaling is dead, so short of figuring out how to 3D stack their wafer scale chips AMD and TSMC are stacking SRAM on the chip scale. I imagine they could accomplish it at the wafer scale. It'll be neat if we can get hundreds of layers in time, like flash. Your analysis seems spot on to me.
Assume you meant Intel, rather than AMD?
Re: Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)
#29> At over 2,500 t/s, Cerebras has set a world record for LLM inference speed on the 400B parameter Llama 4 Maverick model, the largest and most powerful in the Llama 4 family. This is incorrect. The unreleased Llama 4 Behemoth is the largest and most powerful in the Llama 4 family. As for the speed record, it seems important to keep it in context. That comparison is only for performance on 1 query, but it is well kno…
Re: Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)
#30I think it is too risky to build a company around the premise that someone won't soon solve the quadratic scaling issue. Especially, when that company involves creating ASICs. E.g.: https://arxiv.org/abs/2312.00752
Just takes one breakthrough and it's all different. See the recent diffusion style LLMs for example