Live data from Hacker News

Cerebras CS-4

cerebras.ai

131–140 of 281 posts

Re: Cerebras CS-4

#131

Earlier quoted context omitted.

Fable is most definitely nowhere near 10T. The cost to train and infer that would be insane, even by today's standards.

Kimi K3 is a 2.8T model that's available at about 1/4-1/3 the cost of Fable from multiple providers on openrouter. The math doesn't seem wildly off.

The raw margins on proprietary model inference are rumored to be quite high though (they have to successfully defray the entire investment into model training and datacenter capacity for inference, which is massive enough). The API cost you're paying for the model includes that raw margin.

Re: Cerebras CS-4

#132
post #105

Earlier quoted context omitted.

(Where did you see that?) This was also interesting: "CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters." Was it known that there were 10 trillion parameter models in use? I think the frontier providers keep the size of their models carefully hidden.

Mythos/Fable are around 10T: > According to FT, industry estimates say Anthropic's most advanced Mythos 5 has about 8 trillion parameters and Fable 5 about 5 trillion https://www.reuters.com/technology/bytedance-targets-mega-ai... I believe this report has confused Opus (which is known to be around 5T) and Fable. Other reports say 10T. See for example https://eu.36kr.com/en/p/3760679047267075?ref=explainx where Musk…

I'm confused. I thought Mythos 5 and Fable 5 were exactly the same model just with a different security layer in front of it. Could they mean the Mythos 5 Preview?

Re: Cerebras CS-4

#133
post #113

Earlier quoted context omitted.

or the latest qwen3.8 27B doing so well at ~1/100 the size of K3

What about general knowledge you can get out of it before hallucinations start?

Storing general knowledge in VRAM has always been a dumb idea in the first place.

Re: Cerebras CS-4

#134

Earlier quoted context omitted.

Hence why taalas was one of the best strategic acquisitions of the year. I'm honestly baffled they were not acquired by somebody else (sorry AMD).

Taalas will be one of the great disaster investments of the early AI era. It'll be a near total write-down. The absolute worst market time to etch a model to a chip is right now (very rapid iteration). There is no scenario where they can keep up. The Taalas approach will be viewed as comically foolish within just a few years. Cerebras will win in terms of approach. It's 1998: hey, I can drastically speed up your web…

I would still pay ~500 for a chip that runs 10kt/s of a ~100b model on my machine even if the half life is 6months. I m sure my employer would too.

Qwwen3.5 122b was released 6 months ago and is still best in class overall 100-140 B param model.

Re: Cerebras CS-4

#135

Earlier quoted context omitted.

If this is true, it's even more impressive that some of the open weight models that are <3.5T in size, approx 33% of its size, are within a few points of it in the artificial analysis leaderboard.

You want to take a look at the "Scaling Laws" paper, so you can extrapolate from these numbers.

This paper, as well as the Chinchilla one, aged like milk though.

Re: Cerebras CS-4

#136

AMD along with cerebras may probably compete with NVIDIA monopoly in near future. Also, NVIDIA will have competition form multiple companies. Just my prediction.

Like NVIDIA bought Groq, AMD might do well buying Cerebras.

Re: Cerebras CS-4

#137
post #103

Earlier quoted context omitted.

Fable is most definitely nowhere near 10T. The cost to train and infer that would be insane, even by today's standards.

Fable is strongly believed to be around 10T. The most conservative estimate I've seen is 8T. Eg: https://www.reuters.com/technology/bytedance-targets-mega-ai... That reports Mythos as 8T and Fable as 5T, but I think they mean Opus as 5T, which is widely known, eg: https://eu.36kr.com/en/p/3760679047267075?ref=explainx Both Grok and Bytedance are training 10T models.

The fact that Musk claims Opus is 5T to justify why Grok is far behind should be taken with a massive grain of salt given he's a recidivist mythomaniac.

Honestly if Opus is 5T parameters while being matched by the biggest open models that are at least twice smaller, it would mean that the US is already behind China in the AI race, despite a significant edge in compute.

Re: Cerebras CS-4

#138
post #8

> Introducing the all new Cerebras CS-4, a revolutionary rack-scale solution that delivers upto 30x faster inference compared to GPUs, enhanced economics, and a simple path todeploy [sic] hyperscale capacity. Did nobody proofread this?

Well at lest it's written by a human.

Re: Cerebras CS-4

#139
post #93

Cerebras should slowly also move to dgx/ryzen market for a desktop version for masses at affordable price yet providing substantial tokens/second on desktop

Desktop SRAM isn't really viable because it could cost $100K just to load the model.

...so, in the same ballapark as ddr5? :-)

Re: Cerebras CS-4

#140

Earlier quoted context omitted.

Hence why taalas was one of the best strategic acquisitions of the year. I'm honestly baffled they were not acquired by somebody else (sorry AMD).

Taalas will be one of the great disaster investments of the early AI era. It'll be a near total write-down. The absolute worst market time to etch a model to a chip is right now (very rapid iteration). There is no scenario where they can keep up. The Taalas approach will be viewed as comically foolish within just a few years. Cerebras will win in terms of approach. It's 1998: hey, I can drastically speed up your web…

Etched model into a chip? A… mobile chip eventually? Seems prescient.
Post reply on HN