Live data from Hacker News

Cerebras CS-4

cerebras.ai

171–180 of 281 posts

Re: Cerebras CS-4

#171
post #12

Earlier quoted context omitted.

162 kW

I guess we know why there's a fair bit of investment money going into small modular nuclear reactor startups now.

Actually that's mostly just the military funding those, with a few of them having data center partnerships so they can shield themselves from the criticism of what they really are: military contractors.

Re: Cerebras CS-4

#172
post #134

Earlier quoted context omitted.

Taalas will be one of the great disaster investments of the early AI era. It'll be a near total write-down. The absolute worst market time to etch a model to a chip is right now (very rapid iteration). There is no scenario where they can keep up. The Taalas approach will be viewed as comically foolish within just a few years. Cerebras will win in terms of approach. It's 1998: hey, I can drastically speed up your web…

I would still pay ~500 for a chip that runs 10kt/s of a ~100b model on my machine even if the half life is 6months. I m sure my employer would too. Qwwen3.5 122b was released 6 months ago and is still best in class overall 100-140 B param model.

500 what? You’re missing the unit

Re: Cerebras CS-4

#173
post #105

Earlier quoted context omitted.

Mythos/Fable are around 10T: > According to FT, industry estimates say Anthropic's most advanced Mythos 5 has about 8 trillion parameters and Fable 5 about 5 trillion https://www.reuters.com/technology/bytedance-targets-mega-ai... I believe this report has confused Opus (which is known to be around 5T) and Fable. Other reports say 10T. See for example https://eu.36kr.com/en/p/3760679047267075?ref=explainx where Musk…

> I believe this report has confused Opus (which is known to be around 5T) and Fable. 5T for Opus feels quite high though. DeepSeek V4 Pro is a mere 1.6T and often described as a match with Opus in overall quality. Even the largest open models in common use are around 2.8T.

It's not an Opus match.

The difference is very visible in long tail applications. Exactly where you'd expect parameter count to matter.

Re: Cerebras CS-4

#174
post #4
post #3

Earlier quoted context omitted.

This is part of why I think the data center build-out is a bubble. We've barely scratched the surface when it comes to hardware optimization. We'll see exponential improvements in energy efficiency and speed over the next decade. Exponential, not linear. GPUs really aren't that great for AI. They just happen to be the best chips we have in mass production right now for this work load, and it takes time to field new d…

On the plus side, lots of cheap servers to swoop up :)

Look at their power supply, it’s not something you can run in a home lab. Unfortunately most of that will likely go to the bin eventually :(

Re: Cerebras CS-4

#175
post #8

> Introducing the all new Cerebras CS-4, a revolutionary rack-scale solution that delivers upto 30x faster inference compared to GPUs, enhanced economics, and a simple path todeploy [sic] hyperscale capacity. Did nobody proofread this?

Well at lest it's written by a human.

« Make it look like human written »

Re: Cerebras CS-4

#176

That "GPU" comparison is the vaguest i seen so far

True, it's also "per user", somehow, but I think it's a misleading metric. Cerebras chips take the whole wafer? A single TSMC wafer contains 60 to 65 B200s, assuming 70% yields that's 40ish wafers per die. Cerebras cannot redefine wafer economics.

Depends on what you mean, they have more redundancy which means that the yield can be much higher.

Re: Cerebras CS-4

#177

Five years from now, I don't know why anyone will still be using Nvidia for inference. Note that Cerebras is for inference only, not for training. I understand that Cerebras has competition, but this bodes even more poorly for Nvidia for inference. Nvidia may still have a role to play for training, however.

NVIDIA has the best supply chain in the entire game. They are the only ones who can produce at their scale. You really shouldn’t underestimate their position

Re: Cerebras CS-4

#178
post #94

If cerebars is performing well, why didn't its predecessor, server S-3, become the largest API token provider on OpenRouter, surpassing the official model releases?

Cerebras provides high-speed inference at high cost. It's never going to be the cheapest and thus it will probably remain niche.

But that's supply and demand, not technology. Right now a lot more people want their inference than they can supply. as supply catches up in the next 5-10 years, the underlying tech at scale is probably cheaper than GPUs per token produced.
Post reply on HN