Live data from Hacker News

Cerebras CS-4

cerebras.ai

221–230 of 281 posts

Re: Cerebras CS-4

#221

Earlier quoted context omitted.

it only takes ~445 GB300 NVL72 (about $22b) to run ALL of openrouter demand for a year. Microsoft rolled out $32b of DC 2026Q1. imo the issue is that most openrouter demand is inauthentic activity (things that anthropic and openai models will refuse to do like pretend to not be bots when interacting with humans)

I thought your numbers must be wrong. So I plugged 288 trillion tokens/month (OpenRouter's current rate), 500 billion MoE model average, and the math comes out to be around 620 B200 GPUs minimum. So basically, OpenRouter's volume must be absolutely tiny compared to the volume hyperscalers are getting.

For reference, Google serves >3 quadrillion/mo.

[1] https://x.com/ren_stocks/status/2056946641815396718?s=20

Re: Cerebras CS-4

#223

Earlier quoted context omitted.

if you are developing your own hardware, you provide your own stack to avoid lawsuits with nvidia. i don't think it's a technical problem at all, but a legal one. this is probably why zluda was scrapped by AMD and Intel. Nvidia technically bans the creation of CUDA reimplementations in their TOS if i remember correctly

Doesn't Google v Oracle provide protection here? Copying APIs is fair use. If it's patents that are the problem then presumably all these large semiconductor companies have defensive parent portfolios.

I don't know then, amd and Intel decided to get rid of Zluda before that lawsuit ruled that apis are fair use. And now, their cloud customers are ok with using their own stack such as rocm, and Intel got rid of their Garuda ai chips. Those things also happened before the lawsuit was settled

It would be a breach of contract not a copyright issue.

Re: Cerebras CS-4

#224

Earlier quoted context omitted.

pretty sure 10 trillion parameters is now the norm among closed ai labs, given that nvidia also references the same 10 trillion number for their nvl72 racks

Pretty bad efficiency then unless that only applies to Fable class but even then - Kimi K3 is around 3T and does similarly well in most benchmarks.

As others have noted, most benchmarks stress the torso, not the tail.

Re: Cerebras CS-4

#225

Earlier quoted context omitted.

There’s been a spate of Reddit AI bots using all lower case in hopes of evading detection. It’s still incredibly obvious.

I don’t really visit Reddit much these days but would love to see an example.

https://www.reddit.com/r/ModSupport/comments/1tu67kn/influx_...

Re: Cerebras CS-4

#226

God I wish they'd back up all of those claims by offering a subscription of Kimi K3 and GLM 5.3, not some outdated GLM 4.7 instance that they then proceed to call a preview model and say that they'll remove it, leaving users only with GPT-OSS 120B which is nigh useless nowadays: https://support.cerebras.net/articles/9996007307-cerebras-co... and https://www.cerebras.ai/pricing Guess they don't care about regular devs…

Why they should go with Chinese models if they have a line up of gpt models and a very good partnership with someone who lives in the same jurisdiction and not in the country that convinces their citizen that it’s a good idea to go on war with western world ? Just curious ?

Re: Cerebras CS-4

#227
I was hoping to see Cerebras launch something other than GPT-OSS-120b in production this week, especially with GLM4.7 going away.

If they could launch Qwen 27b or Deepseek Flash that would be amazing.

Re: Cerebras CS-4

#228

God I wish they'd back up all of those claims by offering a subscription of Kimi K3 and GLM 5.3, not some outdated GLM 4.7 instance that they then proceed to call a preview model and say that they'll remove it, leaving users only with GPT-OSS 120B which is nigh useless nowadays: https://support.cerebras.net/articles/9996007307-cerebras-co... and https://www.cerebras.ai/pricing Guess they don't care about regular devs…

>Guess they don't care about regular devs atm and are focused only on hardware sales.

Why would they want to target regular devs right now? If they sold to regular devs instead of enterprises, the complaint wouldn't be about model choice, it'd be about how expensive it.

Re: Cerebras CS-4

#229

Earlier quoted context omitted.

Kimi K3 is a 2.8T model that's available at about 1/4-1/3 the cost of Fable from multiple providers on openrouter. The math doesn't seem wildly off.

The raw margins on proprietary model inference are rumored to be quite high though (they have to successfully defray the entire investment into model training and datacenter capacity for inference, which is massive enough). The API cost you're paying for the model includes that raw margin.

The Chinese have similarly high profit margins on Inference via their first party api

Re: Cerebras CS-4

#230
post #163

I think the fun takeaway from this is that GPT 5.4 is probably 45B active parameters and GPT 5.6 Sol is closer to 50B.

You cannot infer this because they only show the tokens per second per user. One way to get a higher number is to have fewer users per chip. I'm pretty sure Cerebras has a confidentiality agreement with OpenAI, and this press release was carefully constructed to avoid leaking details about the model weights. For example, the graph of tokens per second vs. tokens per second per user doesn't have any numbers that would…

They show that CS-4 can't really do batching (or rather it can't properly benefit from it), total throughput barely changes (25%?): https://cdn.sanity.io/images/e4qjo92p/production/6a132331880...

Which I think makes it feasible to approximate activation from CS-4 tokens per second per user.

Post reply on HN