Live data from Hacker News

Cerebras CS-4

cerebras.ai

211–220 of 281 posts

Re: Cerebras CS-4

#211
Impressive that this is an "interim" product, the start of a new line that ought to be continued with the WSE-4 family, where they are supposed to use a 3nm process and, maybe, 3D stacked SRAM. The modular architecture also points towards field upgrades that are badly needed for AI datacenter builders.

Re: Cerebras CS-4

#212

Earlier quoted context omitted.

I guess we know why there's a fair bit of investment money going into small modular nuclear reactor startups now.

And advanced geothermal. Fervo Energy let's us get energy that's not based on burning fossil fuels but is, instead, able to produce energy from the ground.

Remove energy from the ground. I wonder what the consequences may be once we are cooling the underground at several MW/h.

Re: Cerebras CS-4

#213

God I wish they'd back up all of those claims by offering a subscription of Kimi K3 and GLM 5.3, not some outdated GLM 4.7 instance that they then proceed to call a preview model and say that they'll remove it, leaving users only with GPT-OSS 120B which is nigh useless nowadays: https://support.cerebras.net/articles/9996007307-cerebras-co... and https://www.cerebras.ai/pricing Guess they don't care about regular devs…

I don’t think they will until they change the architecture. They don’t have a prefix cache like other providers, or at least don’t have a discount in their billing structure. Each message charges for the whole context window. It’s wildly more expensive for long multi turn scenarios with lots of tool calls (coding). It’s better for short few turn tasks. Edit: I don’t know if they actually have a proper cache. This cou…

Cerebras supports prompt caching and has a doc about it. A fairly standard automatic prefix-based implementation with 5min expiry.

They do not seem to discount cached input for the self-serve Developer tier. Maybe they do for enterprise rate cards?

https://inference-docs.cerebras.ai/capabilities/prompt-cachi...

Re: Cerebras CS-4

#214

God I wish they'd back up all of those claims by offering a subscription of Kimi K3 and GLM 5.3, not some outdated GLM 4.7 instance that they then proceed to call a preview model and say that they'll remove it, leaving users only with GPT-OSS 120B which is nigh useless nowadays: https://support.cerebras.net/articles/9996007307-cerebras-co... and https://www.cerebras.ai/pricing Guess they don't care about regular devs…

Why would they offer a coding subscription and start competing with some of their biggest customers; when they are capacity-bound and companies like OpenAI will take however many wafers that Cerebras sells to them?

OpenAI's Sol ultrafast (powered by Cerebras) is still in preview, presumably because they're overall capacity bound.

Re: Cerebras CS-4

#216
post #209

Earlier quoted context omitted.

The fact that Musk claims Opus is 5T to justify why Grok is far behind should be taken with a massive grain of salt given he's a recidivist mythomaniac. Honestly if Opus is 5T parameters while being matched by the biggest open models that are at least twice smaller, it would mean that the US is already behind China in the AI race, despite a significant edge in compute.

The open models don't really match Opus. For example I regularly do Fable+Opus agentic coding runs over 24 hours without intervention. I think I've had GLM do a run that was a few hours. That's the closest I've had an open model come on that kind of work.

Even if they don't match current-day Opus in everything, they do beat 6 month old Opus, which we have no reason to believe it was smaller than the latest version.

Re: Cerebras CS-4

#217
post #86

Earlier quoted context omitted.

Why would they? What the upside, for them?

Yeah, its not like this js some sort of Open AI company. That'd be ridiculous.

You think their name means releasing competitive details would be good for them?

I’ve got sone bad news about Federal Express.

Re: Cerebras CS-4

#218

I think the fun takeaway from this is that GPT 5.4 is probably 45B active parameters and GPT 5.6 Sol is closer to 50B.

Didn't they say that they can support bigger models now?

Memory capacity on the WSE is the same as before, but access to off-wafer memory is much slower, so the sweet spot is a given fixed balance of memory and compute. They have announced a partnership with AMD in which CPU/GPU hardware is used for part of the workload and the WSE-3 machines are used for inference for specialized smaller models, but I'm not really sure of the details on that.

And there is, of course, the educated guesses about what WSE-4 will be, one being adding a LOT of stacked SRAM or DRAM to tip the balance towards memory (which could also be done by having a few different tile designs with various configurations of compute and memory capacity). I am curious about which way they'll go.

Re: Cerebras CS-4

#219

Earlier quoted context omitted.

You just proved that AI cannot currently do that

It can definitely create a software stack for you if you hold it right, but the software stack supported by a trillion dollar company with decades of expertise, that also uses AI to improve its stack is probably gonna be better.

On the reverse side, there are fundamental limits to the number of ways you can perform certain actions, and agents are both diligent as well as able to swarm. If you have your tests beforehand, there is a chance.

Re: Cerebras CS-4

#220

Earlier quoted context omitted.

What do they do with old ones? Their hardware physically can't run other models right?

The Cerebras hardware is not locked to specific models / model families. Taalas is the company that's etching models into their silicon, locking it to that model forever.

That's pretty sweet. Thanks
Post reply on HN