Live data from Hacker News

Cerebras CS-4

cerebras.ai

41–50 of 281 posts

Re: Cerebras CS-4

#42

If cerebars is performing well, why didn't its predecessor, server S-3, become the largest API token provider on OpenRouter, surpassing the official model releases?

it only takes ~445 GB300 NVL72 (about $22b) to run ALL of openrouter demand for a year. Microsoft rolled out $32b of DC 2026Q1.

imo the issue is that most openrouter demand is inauthentic activity (things that anthropic and openai models will refuse to do like pretend to not be bots when interacting with humans)

Re: Cerebras CS-4

#43

If cerebars is performing well, why didn't its predecessor, server S-3, become the largest API token provider on OpenRouter, surpassing the official model releases?

Without having any inside information, one possible theory:

All or a vast majority of of the cerebras manufacturing capacity was going to a few companies that aren't publicly available inference providers on openrouter, for their own internal use.

or

The asking price of the S-3, no matter how speedy it might be, for small/medium size customers made it economically prohibitive to purchase and use to sell public inference vs. buying more common nvidia b200 or whatever.

Re: Cerebras CS-4

#44
post #14

Earlier quoted context omitted.

Congratulations! You have just realized that the AI data center build out is a total scam, built on both the insurmountable trillions of debt, and the assumption that only GPUs are all we need to continue scaling. There exist other AI accelerators (TPUs, ASICs) that perfectly exceed the throughput that LLMs need to scale as well. But the true solution is more software optimizations. There's a tiny handful of them but…

I wonder what this looks like in 5 years... Will there be a massive push to repurpose these giant boxes into housing? Will they get turned back into the farm land from where they came? When a data center goes bust, what happens to the parts left behind?

I'd think the infrastructure would tend towards factories, smelters, and so on. Industrial things that have reasonably high power demands, can use the building, and don't care about the lack of windows.

They're typically not built where you want housing, and the buildings are distinctly the wrong shape.

If you can't use the power infrastructure profitably my next thought would be warehousing.

But also... we've seen a pretty continually increasing demand for compute. Even if AI busts a bit (or becomes a bit more efficient) I bet most data centres stay data centres, just less profitable ones.

Re: Cerebras CS-4

#45

If cerebars is performing well, why didn't its predecessor, server S-3, become the largest API token provider on OpenRouter, surpassing the official model releases?

If you're willing to pay a significant premium for latency, why use openrouter? And anyway Cerebras only supported a few specific models.

Re: Cerebras CS-4

#46

> CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters Oops did they just out GPT-5.6 sol’s parameter count?

Sol is supposed to be 5T according to rumour. The imminent Astra is allegedly 10

Re: Cerebras CS-4

#47
post #40

Earlier quoted context omitted.

(Where did you see that?) This was also interesting: "CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters." Was it known that there were 10 trillion parameter models in use? I think the frontier providers keep the size of their models carefully hidden.

It's rumored fable is around that 10T number

If this is true, it's even more impressive that some of the open weight models that are <3.5T in size, approx 33% of its size, are within a few points of it in the artificial analysis leaderboard.

Re: Cerebras CS-4

#48
post #24

Earlier quoted context omitted.

I mean the product is a server rack and while there's no advertised price I would assume it's six figures. So yes, an enterprise product. But even an enterprise is going to care about the difference between "we can run the model we want with support from the manufacturer" and "we have to purchase the product, and then spend another 6 figure sum having developers port a recent model to the product to use it".

I feel like 6-figures would be the clearance price on it...

6 figures is a single mid range Xeon or Epyc server these days.

Re: Cerebras CS-4

#50
post #30

KV caching status? What's the point of 1000tok/s if you have to do prefill on every agentic turn which at 100k depth would make it 1.5 min latency every turn?

Information about RAM type/size and connection topology of the RAM to be used for context cache seems to be conspicuously absent from the slick looking marketing materials.

There's a few more details at the bottom of this page: https://www.cerebras.ai/blog/introducing-cerebras-cs-4

44GB on-chip-sram * 3 chips. Per chip: 43.2 PB/s memory access + 53.5 PB/s on-chip fabric bandwidth + 2.4 Tbits/s "IO" bandwidth (I think that means their RoCE v2 RDMA over Ethernet interface).

I suspect there might be a certain amount of customization for how much RAM they attach when you order it.

Post reply on HN