Live data from Hacker News

Cerebras CS-4

cerebras.ai

21–30 of 281 posts

Re: Cerebras CS-4

#21
post #16

Is it just me or is it bizarre that they're advertising old open-weight models. GLM 4.7 (December 2025) not 5 (Feb) 5.1 (April) or 5.2 (June). 5.3 (4 days ago) is, to be fair, not open weights yet... but there's a lot since 4.7. Kimi K2.7 (April) not K2.7-code (June) or K3 (July). Gemma 4 (April), Llama (April), and gpt-oss (August 2025) are up to date, but old (for models). Meanwhile the closed source GPT 5.6 sol is…

I think they run whatever models they get paid to run. But mostly from enterprise. They are clearly not interested in consumer dollars.

Re: Cerebras CS-4

#22
post #8

> Introducing the all new Cerebras CS-4, a revolutionary rack-scale solution that delivers upto 30x faster inference compared to GPUs, enhanced economics, and a simple path todeploy [sic] hyperscale capacity. Did nobody proofread this?

Sometimes I wonder if mistakes are now used to indicate the possibility that a human actually wrote it.

Re: Cerebras CS-4

#23
post #17
post #12

Earlier quoted context omitted.

162 kW

I presume per rack? Can you imagine something radiating that much energy into a space in your home?

I guess because I have actually set foot in a data center I don't imagine literally every product in my home.

Re: Cerebras CS-4

#24
post #21
post #16

Is it just me or is it bizarre that they're advertising old open-weight models. GLM 4.7 (December 2025) not 5 (Feb) 5.1 (April) or 5.2 (June). 5.3 (4 days ago) is, to be fair, not open weights yet... but there's a lot since 4.7. Kimi K2.7 (April) not K2.7-code (June) or K3 (July). Gemma 4 (April), Llama (April), and gpt-oss (August 2025) are up to date, but old (for models). Meanwhile the closed source GPT 5.6 sol is…

I think they run whatever models they get paid to run. But mostly from enterprise. They are clearly not interested in consumer dollars.

I mean the product is a server rack and while there's no advertised price I would assume it's six figures. So yes, an enterprise product.

But even an enterprise is going to care about the difference between "we can run the model we want with support from the manufacturer" and "we have to purchase the product, and then spend another 6 figure sum having developers port a recent model to the product to use it".

Re: Cerebras CS-4

#26
post #24
post #21

Earlier quoted context omitted.

I think they run whatever models they get paid to run. But mostly from enterprise. They are clearly not interested in consumer dollars.

I mean the product is a server rack and while there's no advertised price I would assume it's six figures. So yes, an enterprise product. But even an enterprise is going to care about the difference between "we can run the model we want with support from the manufacturer" and "we have to purchase the product, and then spend another 6 figure sum having developers port a recent model to the product to use it".

I feel like 6-figures would be the clearance price on it...

Re: Cerebras CS-4

#27
post #14

Just a reminder for everyone that we are only several years and 3 or 4 iterations into hardware being optimized for LLMs. We should all expect orders of magnitude improvement in speed and/or cost over the next 5 years. Then we can have fun conversations about "unlimited" "intelligence" and about what the price wars and profit margins of consumer AI products are when your average ChatGPT user costs the company $0.10 p…

Congratulations! You have just realized that the AI data center build out is a total scam, built on both the insurmountable trillions of debt, and the assumption that only GPUs are all we need to continue scaling. There exist other AI accelerators (TPUs, ASICs) that perfectly exceed the throughput that LLMs need to scale as well. But the true solution is more software optimizations. There's a tiny handful of them but…

TPUs and ASICs run in data centers too. Your argument only holds true if there's some satisfied limit to demand for inference. If not, data centers will continue to spring up to host more and more agents. Even if agents were running on hardware and software as efficient as the human brain, its conceivable we want trillions of them running at any given time which would require data center scale.

Re: Cerebras CS-4

#30
KV caching status?

What's the point of 1000tok/s if you have to do prefill on every agentic turn which at 100k depth would make it 1.5 min latency every turn?

Post reply on HN