Live data from Hacker News

Cerebras CS-4

cerebras.ai

121–130 of 281 posts

Re: Cerebras CS-4

#121

I think the fun takeaway from this is that GPT 5.4 is probably 45B active parameters and GPT 5.6 Sol is closer to 50B.

(Where did you see that?) This was also interesting: "CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters." Was it known that there were 10 trillion parameter models in use? I think the frontier providers keep the size of their models carefully hidden.

You don't really need to train a 10T model to test cerebras against a 10T model. You can feed it an untrained (randomly initialized) model and benchmark it. Result will be gibberish but performance the same.

Re: Cerebras CS-4

#122
post #103

Earlier quoted context omitted.

Fable is most definitely nowhere near 10T. The cost to train and infer that would be insane, even by today's standards.

Fable is strongly believed to be around 10T. The most conservative estimate I've seen is 8T. Eg: https://www.reuters.com/technology/bytedance-targets-mega-ai... That reports Mythos as 8T and Fable as 5T, but I think they mean Opus as 5T, which is widely known, eg: https://eu.36kr.com/en/p/3760679047267075?ref=explainx Both Grok and Bytedance are training 10T models.

Wasn't Opus ~1.5T and Fable is about twice that?

Re: Cerebras CS-4

#123

Earlier quoted context omitted.

Hence why taalas was one of the best strategic acquisitions of the year. I'm honestly baffled they were not acquired by somebody else (sorry AMD).

Taalas will be one of the great disaster investments of the early AI era. It'll be a near total write-down. The absolute worst market time to etch a model to a chip is right now (very rapid iteration). There is no scenario where they can keep up. The Taalas approach will be viewed as comically foolish within just a few years. Cerebras will win in terms of approach. It's 1998: hey, I can drastically speed up your web…

> The absolute worst market time to etch a model to a chip is right now

Slightly disagree. It really depends on the price-point at which they can do that etching. ~1k usd / ~30B model in a hdd-sized case that fits on your desk? I'd buy one right now, even knowing that I'm "stuck" with whatever model of the day is.

Re: Cerebras CS-4

#124
post #40

Earlier quoted context omitted.

It's rumored fable is around that 10T number

Fable is most definitely nowhere near 10T. The cost to train and infer that would be insane, even by today's standards.

Kimi K3 is a 2.8T model that's available at about 1/4-1/3 the cost of Fable from multiple providers on openrouter. The math doesn't seem wildly off.

Re: Cerebras CS-4

#125

Earlier quoted context omitted.

Taalas will be one of the great disaster investments of the early AI era. It'll be a near total write-down. The absolute worst market time to etch a model to a chip is right now (very rapid iteration). There is no scenario where they can keep up. The Taalas approach will be viewed as comically foolish within just a few years. Cerebras will win in terms of approach. It's 1998: hey, I can drastically speed up your web…

It's 2026: let's etch nginx into silicon and get 10,000,000 rps at a cost of 0.1 US/day. Yes, please!

Most people probably don't care about nginx performance. It shouldn't be your bottleneck unless you serve massive amounts of static data.

Re: Cerebras CS-4

#126

Earlier quoted context omitted.

Is CUDA still a moat? Are we not at the point where frontier models can reimplement software stacks, given you throw enough tokens at the problem.

if you are developing your own hardware, you provide your own stack to avoid lawsuits with nvidia. i don't think it's a technical problem at all, but a legal one. this is probably why zluda was scrapped by AMD and Intel. Nvidia technically bans the creation of CUDA reimplementations in their TOS if i remember correctly

Doesn't Google v Oracle provide protection here? Copying APIs is fair use.

If it's patents that are the problem then presumably all these large semiconductor companies have defensive parent portfolios.

Re: Cerebras CS-4

#127

Earlier quoted context omitted.

Taalas will be one of the great disaster investments of the early AI era. It'll be a near total write-down. The absolute worst market time to etch a model to a chip is right now (very rapid iteration). There is no scenario where they can keep up. The Taalas approach will be viewed as comically foolish within just a few years. Cerebras will win in terms of approach. It's 1998: hey, I can drastically speed up your web…

> The absolute worst market time to etch a model to a chip is right now Slightly disagree. It really depends on the price-point at which they can do that etching. ~1k usd / ~30B model in a hdd-sized case that fits on your desk? I'd buy one right now, even knowing that I'm "stuck" with whatever model of the day is.

Time to market also matters a ton. If they can start shipping chips <1 month after the weights drop that's much more compelling than if it's a 6+ month development pipeline.

Re: Cerebras CS-4

#128
post #34
post #12

Earlier quoted context omitted.

162 kW

God, I was going to ask if this could be deployed in a standard existing datacenter, but I guess that answers that question.

So the answer is yes? Putting in a few of those racks for special tasks shouldn't break the power assumptions of a data center.

Re: Cerebras CS-4

#129

Earlier quoted context omitted.

Is CUDA still a moat? Are we not at the point where frontier models can reimplement software stacks, given you throw enough tokens at the problem.

if you are developing your own hardware, you provide your own stack to avoid lawsuits with nvidia. i don't think it's a technical problem at all, but a legal one. this is probably why zluda was scrapped by AMD and Intel. Nvidia technically bans the creation of CUDA reimplementations in their TOS if i remember correctly

[flagged]

Re: Cerebras CS-4

#130
post #105

Earlier quoted context omitted.

(Where did you see that?) This was also interesting: "CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters." Was it known that there were 10 trillion parameter models in use? I think the frontier providers keep the size of their models carefully hidden.

Mythos/Fable are around 10T: > According to FT, industry estimates say Anthropic's most advanced Mythos 5 has about 8 trillion parameters and Fable 5 about 5 trillion https://www.reuters.com/technology/bytedance-targets-mega-ai... I believe this report has confused Opus (which is known to be around 5T) and Fable. Other reports say 10T. See for example https://eu.36kr.com/en/p/3760679047267075?ref=explainx where Musk…

> I believe this report has confused Opus (which is known to be around 5T) and Fable.

5T for Opus feels quite high though. DeepSeek V4 Pro is a mere 1.6T and often described as a match with Opus in overall quality. Even the largest open models in common use are around 2.8T.

Post reply on HN