Live data from Hacker News

Cerebras CS-4

cerebras.ai

71–80 of 281 posts

Re: Cerebras CS-4

#71
post #24
post #21

Earlier quoted context omitted.

I think they run whatever models they get paid to run. But mostly from enterprise. They are clearly not interested in consumer dollars.

I mean the product is a server rack and while there's no advertised price I would assume it's six figures. So yes, an enterprise product. But even an enterprise is going to care about the difference between "we can run the model we want with support from the manufacturer" and "we have to purchase the product, and then spend another 6 figure sum having developers port a recent model to the product to use it".

You’re at least an order or magnitude under… likely two.

A single AI server with a mere 8 GPUs from Nvidia is already mid 6 digits. A rack system from Nvidia is mid 7 digits.

There’s some info out there that suggests the CS1 had an 8 digits price tag, so it wouldn’t be surprising to see that here.

Re: Cerebras CS-4

#72
post #40

Earlier quoted context omitted.

It's rumored fable is around that 10T number

If this is true, it's even more impressive that some of the open weight models that are <3.5T in size, approx 33% of its size, are within a few points of it in the artificial analysis leaderboard.

GLM 5.3 is "only" 753B parameters. Much much smaller.

Re: Cerebras CS-4

#73
post #24

Earlier quoted context omitted.

I mean the product is a server rack and while there's no advertised price I would assume it's six figures. So yes, an enterprise product. But even an enterprise is going to care about the difference between "we can run the model we want with support from the manufacturer" and "we have to purchase the product, and then spend another 6 figure sum having developers port a recent model to the product to use it".

You’re at least an order or magnitude under… likely two. A single AI server with a mere 8 GPUs from Nvidia is already mid 6 digits. A rack system from Nvidia is mid 7 digits. There’s some info out there that suggests the CS1 had an 8 digits price tag, so it wouldn’t be surprising to see that here.

Yeah, did some googling after WarmWash's comment and I concur.

Re: Cerebras CS-4

#74
post #46

> CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters Oops did they just out GPT-5.6 sol’s parameter count?

Sol is supposed to be 5T according to rumour. The imminent Astra is allegedly 10

Rumors and allegations aren't worth much. Why don't they just tell us mere mortals?

Re: Cerebras CS-4

#75

If cerebars is performing well, why didn't its predecessor, server S-3, become the largest API token provider on OpenRouter, surpassing the official model releases?

it only takes ~445 GB300 NVL72 (about $22b) to run ALL of openrouter demand for a year. Microsoft rolled out $32b of DC 2026Q1. imo the issue is that most openrouter demand is inauthentic activity (things that anthropic and openai models will refuse to do like pretend to not be bots when interacting with humans)

I was curious so I looked it up: looks like a GB300 NVL72 is about $4M. So $22B would buy you 5500 such racks, no?

Re: Cerebras CS-4

#76

Earlier quoted context omitted.

Sometimes I wonder if mistakes are now used to indicate the possibility that a human actually wrote it.

There’s been a spate of Reddit AI bots using all lower case in hopes of evading detection. It’s still incredibly obvious.

I don’t really visit Reddit much these days but would love to see an example.

Re: Cerebras CS-4

#77
post #4

Earlier quoted context omitted.

On the plus side, lots of cheap servers to swoop up :)

But power hungry. In that 5+ year timeline, the compute per watt could change by three orders of magnitude. GPUs are to LLMs what CPUs are to gaming — not a good fit.

A cursory estimate courtesy of ChatGPT suggests that there is a grand total of one order of magnitude or less of power efficiency improvement available compared to current Blackwell if the entire system’s power consumption outside the ALUs went all the way to zero.

If you want three orders of magnitude improvement, you probably need to find two of those orders of magnitude somewhere else: process improvements, different ALU design, model architecture changes, etc.

Re: Cerebras CS-4

#78

If cerebars is performing well, why didn't its predecessor, server S-3, become the largest API token provider on OpenRouter, surpassing the official model releases?

it only takes ~445 GB300 NVL72 (about $22b) to run ALL of openrouter demand for a year. Microsoft rolled out $32b of DC 2026Q1. imo the issue is that most openrouter demand is inauthentic activity (things that anthropic and openai models will refuse to do like pretend to not be bots when interacting with humans)

It is worth mentioning, the OpenRouter demand isn't static though. It has increased week on week since early 2026.

Re: Cerebras CS-4

#79
post #3

Just a reminder for everyone that we are only several years and 3 or 4 iterations into hardware being optimized for LLMs. We should all expect orders of magnitude improvement in speed and/or cost over the next 5 years. Then we can have fun conversations about "unlimited" "intelligence" and about what the price wars and profit margins of consumer AI products are when your average ChatGPT user costs the company $0.10 p…

This is part of why I think the data center build-out is a bubble. We've barely scratched the surface when it comes to hardware optimization. We'll see exponential improvements in energy efficiency and speed over the next decade. Exponential, not linear. GPUs really aren't that great for AI. They just happen to be the best chips we have in mass production right now for this work load, and it takes time to field new d…

By the time these gigawatt datacenters are done being built the hardware will be so far behind state of the art they may be mostly useless.

Re: Cerebras CS-4

#80

AMD along with cerebras may probably compete with NVIDIA monopoly in near future. Also, NVIDIA will have competition form multiple companies. Just my prediction.

Maybe, but GPU is just one aspect of NVIDIA's dominance. If you are buying Vera Rubin GPUs, you're getting an NVL72 rack, which is only one of several racks that you're probably buying. You'll also need your NVIDIA racks with NVIDIA networking & storage gear, too. At the end of the day, they're "vertically integrated" for your accelerated computing data center (e.g. the "AI Factory"). This doesn't even count the software layer, where CUDA + CUDA-X (not to mention the software for all the sysadmin pieces) has a huge first mover advantage over anyone else.
Post reply on HN