Live data from Hacker News

Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)

cerebras.ai

11–20 of 100 posts

Re: Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)

#11
post #3

> At over 2,500 t/s, Cerebras has set a world record for LLM inference speed on the 400B parameter Llama 4 Maverick model, the largest and most powerful in the Llama 4 family. This is incorrect. The unreleased Llama 4 Behemoth is the largest and most powerful in the Llama 4 family. As for the speed record, it seems important to keep it in context. That comparison is only for performance on 1 query, but it is well kno…

> Also, Cerebras is the company that not only was saying that their hardware was not useful for inference until some time last year, but even partnered with Qualcomm with the claim that Qualcomm’s accelerators had a 10x price performance improvement over their things

Mistral says they run Le Chat on Cerebras

Re: Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)

#12
post #3

> At over 2,500 t/s, Cerebras has set a world record for LLM inference speed on the 400B parameter Llama 4 Maverick model, the largest and most powerful in the Llama 4 family. This is incorrect. The unreleased Llama 4 Behemoth is the largest and most powerful in the Llama 4 family. As for the speed record, it seems important to keep it in context. That comparison is only for performance on 1 query, but it is well kno…

>Each one costs ~$2 million, so that is $40 million.

Pricing for exotic hardware that is not manufactured at scale is quite meaningless. They are selling tokens over an API. The token pricing is competitive with other token APIs.

Re: Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)

#14
post #4
post #2

Investors list include Altman and Ilya https://www.cerebras.ai/company

Their CEO is a felon who plead guilty to accounting fraud: https://milled.com/theinformation/cerebras-ceos-past-felony-... Experienced investors will not touch them: https://www.nbclosangeles.com/news/business/money-report/cer... I estimated last year that they can only produce about 300 chips per year and that is unlikely to change because there are far bigger customers for TSMC that are ahead of them in priority fo…

While the CEO stuff is a problem, I don't think the other stuff matters.

Per chip area WSE-3 is only a little bit more expensive than H200. While you may need several WSE-3s to load the model, if you have enough demand that you are running the WSE-3 at full speed you will not be using more area in the WSE-3. In fact, the WSE-3 may be more efficient, since it won't be loading and unloading things from large memories.

The only effect is that the WSE-3s will have a minimum demand before they make sense, whereas an H200 will make sense even with little demand.

Re: Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)

#15
post #3

> At over 2,500 t/s, Cerebras has set a world record for LLM inference speed on the 400B parameter Llama 4 Maverick model, the largest and most powerful in the Llama 4 family. This is incorrect. The unreleased Llama 4 Behemoth is the largest and most powerful in the Llama 4 family. As for the speed record, it seems important to keep it in context. That comparison is only for performance on 1 query, but it is well kno…

>Their hardware does inference with FP16, so they need ~20 of their CSE-3 chips to run this model.

Care to explain? I don't see it.

Re: Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)

#16
post #4
post #2

Investors list include Altman and Ilya https://www.cerebras.ai/company

Their CEO is a felon who plead guilty to accounting fraud: https://milled.com/theinformation/cerebras-ceos-past-felony-... Experienced investors will not touch them: https://www.nbclosangeles.com/news/business/money-report/cer... I estimated last year that they can only produce about 300 chips per year and that is unlikely to change because there are far bigger customers for TSMC that are ahead of them in priority fo…

>Their CEO is a felon who plead guilty to accounting fraud [...]

Whoa, I didn't know that.

I know he's very close to another guy I know first hand to be a criminal. I won't write the name here for obvious reasons, also not my fight to fight.

I always thought it was a bit weird of them to hang around because I never got that vibe from Feldman, but ... now I came to know about this, 2nd strike I guess ...

Re: Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)

#17
post #8

Maybe one day they’ll have an actual api that you can pay per token. Right now it’s the standard “talk to us” if you want to use it.

Huh? Just make an account, get your API key, and try out the free tier.. works for me.

https://cloud.cerebras.ai

Re: Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)

#20
post #4

Earlier quoted context omitted.

Their CEO is a felon who plead guilty to accounting fraud: https://milled.com/theinformation/cerebras-ceos-past-felony-... Experienced investors will not touch them: https://www.nbclosangeles.com/news/business/money-report/cer... I estimated last year that they can only produce about 300 chips per year and that is unlikely to change because there are far bigger customers for TSMC that are ahead of them in priority fo…

>Their CEO is a felon who plead guilty to accounting fraud [...] Whoa, I didn't know that. I know he's very close to another guy I know first hand to be a criminal. I won't write the name here for obvious reasons, also not my fight to fight. I always thought it was a bit weird of them to hang around because I never got that vibe from Feldman, but ... now I came to know about this, 2nd strike I guess ...

CNBC lists several other red flags (one customer generating >80% of revenue, non-top-tier investment bank/auditor).

see https://www.cnbc.com/2024/10/11/cerebras-ipo-has-too-much-ha...

IPO was supposed to happen in autumn 2024.

Post reply on HN