I think the fun takeaway from this is that GPT 5.4 is probably 45B active parameters and GPT 5.6 Sol is closer to 50B.
(Where did you see that?) This was also interesting: "CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters." Was it known that there were 10 trillion parameter models in use? I think the frontier providers keep the size of their models carefully hidden.
Cerebras CS-4
121–130 of 281 posts
Re: Cerebras CS-4
#122Earlier quoted context omitted.
Fable is most definitely nowhere near 10T. The cost to train and infer that would be insane, even by today's standards.
Fable is strongly believed to be around 10T. The most conservative estimate I've seen is 8T. Eg: https://www.reuters.com/technology/bytedance-targets-mega-ai... That reports Mythos as 8T and Fable as 5T, but I think they mean Opus as 5T, which is widely known, eg: https://eu.36kr.com/en/p/3760679047267075?ref=explainx Both Grok and Bytedance are training 10T models.
Re: Cerebras CS-4
#123Earlier quoted context omitted.
Hence why taalas was one of the best strategic acquisitions of the year. I'm honestly baffled they were not acquired by somebody else (sorry AMD).
Taalas will be one of the great disaster investments of the early AI era. It'll be a near total write-down. The absolute worst market time to etch a model to a chip is right now (very rapid iteration). There is no scenario where they can keep up. The Taalas approach will be viewed as comically foolish within just a few years. Cerebras will win in terms of approach. It's 1998: hey, I can drastically speed up your web…
Slightly disagree. It really depends on the price-point at which they can do that etching. ~1k usd / ~30B model in a hdd-sized case that fits on your desk? I'd buy one right now, even knowing that I'm "stuck" with whatever model of the day is.
Re: Cerebras CS-4
#124Earlier quoted context omitted.
It's rumored fable is around that 10T number
Fable is most definitely nowhere near 10T. The cost to train and infer that would be insane, even by today's standards.
Re: Cerebras CS-4
#125Earlier quoted context omitted.
Taalas will be one of the great disaster investments of the early AI era. It'll be a near total write-down. The absolute worst market time to etch a model to a chip is right now (very rapid iteration). There is no scenario where they can keep up. The Taalas approach will be viewed as comically foolish within just a few years. Cerebras will win in terms of approach. It's 1998: hey, I can drastically speed up your web…
It's 2026: let's etch nginx into silicon and get 10,000,000 rps at a cost of 0.1 US/day. Yes, please!
Re: Cerebras CS-4
#126Earlier quoted context omitted.
Is CUDA still a moat? Are we not at the point where frontier models can reimplement software stacks, given you throw enough tokens at the problem.
if you are developing your own hardware, you provide your own stack to avoid lawsuits with nvidia. i don't think it's a technical problem at all, but a legal one. this is probably why zluda was scrapped by AMD and Intel. Nvidia technically bans the creation of CUDA reimplementations in their TOS if i remember correctly
If it's patents that are the problem then presumably all these large semiconductor companies have defensive parent portfolios.
Re: Cerebras CS-4
#127Earlier quoted context omitted.
Taalas will be one of the great disaster investments of the early AI era. It'll be a near total write-down. The absolute worst market time to etch a model to a chip is right now (very rapid iteration). There is no scenario where they can keep up. The Taalas approach will be viewed as comically foolish within just a few years. Cerebras will win in terms of approach. It's 1998: hey, I can drastically speed up your web…
> The absolute worst market time to etch a model to a chip is right now Slightly disagree. It really depends on the price-point at which they can do that etching. ~1k usd / ~30B model in a hdd-sized case that fits on your desk? I'd buy one right now, even knowing that I'm "stuck" with whatever model of the day is.
Re: Cerebras CS-4
#128Earlier quoted context omitted.
162 kW
God, I was going to ask if this could be deployed in a standard existing datacenter, but I guess that answers that question.
Re: Cerebras CS-4
#129Earlier quoted context omitted.
Is CUDA still a moat? Are we not at the point where frontier models can reimplement software stacks, given you throw enough tokens at the problem.
if you are developing your own hardware, you provide your own stack to avoid lawsuits with nvidia. i don't think it's a technical problem at all, but a legal one. this is probably why zluda was scrapped by AMD and Intel. Nvidia technically bans the creation of CUDA reimplementations in their TOS if i remember correctly
Re: Cerebras CS-4
#130Earlier quoted context omitted.
(Where did you see that?) This was also interesting: "CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters." Was it known that there were 10 trillion parameter models in use? I think the frontier providers keep the size of their models carefully hidden.
Mythos/Fable are around 10T: > According to FT, industry estimates say Anthropic's most advanced Mythos 5 has about 8 trillion parameters and Fable 5 about 5 trillion https://www.reuters.com/technology/bytedance-targets-mega-ai... I believe this report has confused Opus (which is known to be around 5T) and Fable. Other reports say 10T. See for example https://eu.36kr.com/en/p/3760679047267075?ref=explainx where Musk…
5T for Opus feels quite high though. DeepSeek V4 Pro is a mere 1.6T and often described as a match with Opus in overall quality. Even the largest open models in common use are around 2.8T.