Live data from Hacker News

Cerebras CS-4

cerebras.ai

31–40 of 281 posts

Re: Cerebras CS-4

#31
post #8

> Introducing the all new Cerebras CS-4, a revolutionary rack-scale solution that delivers upto 30x faster inference compared to GPUs, enhanced economics, and a simple path todeploy [sic] hyperscale capacity. Did nobody proofread this?

Sometimes I wonder if mistakes are now used to indicate the possibility that a human actually wrote it.

There’s been a spate of Reddit AI bots using all lower case in hopes of evading detection.

It’s still incredibly obvious.

Re: Cerebras CS-4

#32

> CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters Oops did they just out GPT-5.6 sol’s parameter count?

I mean we kinda know the frontier models are multi trillion parameter models. The only open weights that are close to the frontier are that size too

Re: Cerebras CS-4

#33
post #14

Just a reminder for everyone that we are only several years and 3 or 4 iterations into hardware being optimized for LLMs. We should all expect orders of magnitude improvement in speed and/or cost over the next 5 years. Then we can have fun conversations about "unlimited" "intelligence" and about what the price wars and profit margins of consumer AI products are when your average ChatGPT user costs the company $0.10 p…

Congratulations! You have just realized that the AI data center build out is a total scam, built on both the insurmountable trillions of debt, and the assumption that only GPUs are all we need to continue scaling. There exist other AI accelerators (TPUs, ASICs) that perfectly exceed the throughput that LLMs need to scale as well. But the true solution is more software optimizations. There's a tiny handful of them but…

I wonder what this looks like in 5 years... Will there be a massive push to repurpose these giant boxes into housing? Will they get turned back into the farm land from where they came? When a data center goes bust, what happens to the parts left behind?

Re: Cerebras CS-4

#35
post #17
post #12

Earlier quoted context omitted.

162 kW

I presume per rack? Can you imagine something radiating that much energy into a space in your home?

It's mandatory liquid cooling, so it's meant to be attached to a specialized liquid cooling loop that gets the heat outside the building.

This is far beyond the practical maximums of like 10 to 15kW per 44U cabinet front to rear air cooling for 'regular' rackmount server stuff.

Re: Cerebras CS-4

#37

I think the fun takeaway from this is that GPT 5.4 is probably 45B active parameters and GPT 5.6 Sol is closer to 50B.

(Where did you see that?)

This was also interesting: "CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters." Was it known that there were 10 trillion parameter models in use?

I think the frontier providers keep the size of their models carefully hidden.

Re: Cerebras CS-4

#38
post #30

KV caching status? What's the point of 1000tok/s if you have to do prefill on every agentic turn which at 100k depth would make it 1.5 min latency every turn?

Information about RAM type/size and connection topology of the RAM to be used for context cache seems to be conspicuously absent from the slick looking marketing materials.

Re: Cerebras CS-4

#40

I think the fun takeaway from this is that GPT 5.4 is probably 45B active parameters and GPT 5.6 Sol is closer to 50B.

(Where did you see that?) This was also interesting: "CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters." Was it known that there were 10 trillion parameter models in use? I think the frontier providers keep the size of their models carefully hidden.

It's rumored fable is around that 10T number
Post reply on HN