Earlier quoted context omitted.
162 kW
I guess we know why there's a fair bit of investment money going into small modular nuclear reactor startups now.
Cerebras CS-4
171–180 of 281 posts
Re: Cerebras CS-4
#172Earlier quoted context omitted.
Taalas will be one of the great disaster investments of the early AI era. It'll be a near total write-down. The absolute worst market time to etch a model to a chip is right now (very rapid iteration). There is no scenario where they can keep up. The Taalas approach will be viewed as comically foolish within just a few years. Cerebras will win in terms of approach. It's 1998: hey, I can drastically speed up your web…
I would still pay ~500 for a chip that runs 10kt/s of a ~100b model on my machine even if the half life is 6months. I m sure my employer would too. Qwwen3.5 122b was released 6 months ago and is still best in class overall 100-140 B param model.
Re: Cerebras CS-4
#173Earlier quoted context omitted.
Mythos/Fable are around 10T: > According to FT, industry estimates say Anthropic's most advanced Mythos 5 has about 8 trillion parameters and Fable 5 about 5 trillion https://www.reuters.com/technology/bytedance-targets-mega-ai... I believe this report has confused Opus (which is known to be around 5T) and Fable. Other reports say 10T. See for example https://eu.36kr.com/en/p/3760679047267075?ref=explainx where Musk…
> I believe this report has confused Opus (which is known to be around 5T) and Fable. 5T for Opus feels quite high though. DeepSeek V4 Pro is a mere 1.6T and often described as a match with Opus in overall quality. Even the largest open models in common use are around 2.8T.
The difference is very visible in long tail applications. Exactly where you'd expect parameter count to matter.
Re: Cerebras CS-4
#174Earlier quoted context omitted.
This is part of why I think the data center build-out is a bubble. We've barely scratched the surface when it comes to hardware optimization. We'll see exponential improvements in energy efficiency and speed over the next decade. Exponential, not linear. GPUs really aren't that great for AI. They just happen to be the best chips we have in mass production right now for this work load, and it takes time to field new d…
On the plus side, lots of cheap servers to swoop up :)
Re: Cerebras CS-4
#175> Introducing the all new Cerebras CS-4, a revolutionary rack-scale solution that delivers upto 30x faster inference compared to GPUs, enhanced economics, and a simple path todeploy [sic] hyperscale capacity. Did nobody proofread this?
Well at lest it's written by a human.
Re: Cerebras CS-4
#176That "GPU" comparison is the vaguest i seen so far
True, it's also "per user", somehow, but I think it's a misleading metric. Cerebras chips take the whole wafer? A single TSMC wafer contains 60 to 65 B200s, assuming 70% yields that's 40ish wafers per die. Cerebras cannot redefine wafer economics.
Re: Cerebras CS-4
#177Five years from now, I don't know why anyone will still be using Nvidia for inference. Note that Cerebras is for inference only, not for training. I understand that Cerebras has competition, but this bodes even more poorly for Nvidia for inference. Nvidia may still have a role to play for training, however.
Re: Cerebras CS-4
#178If cerebars is performing well, why didn't its predecessor, server S-3, become the largest API token provider on OpenRouter, surpassing the official model releases?
Cerebras provides high-speed inference at high cost. It's never going to be the cheapest and thus it will probably remain niche.
Re: Cerebras CS-4
#179What's the sticker price? If I have 20 million in the bank can I just like buy one or what