Live data from Hacker News

Cerebras CS-4

cerebras.ai

161–170 of 281 posts

Re: Cerebras CS-4

#161
post #3

Just a reminder for everyone that we are only several years and 3 or 4 iterations into hardware being optimized for LLMs. We should all expect orders of magnitude improvement in speed and/or cost over the next 5 years. Then we can have fun conversations about "unlimited" "intelligence" and about what the price wars and profit margins of consumer AI products are when your average ChatGPT user costs the company $0.10 p…

This is part of why I think the data center build-out is a bubble. We've barely scratched the surface when it comes to hardware optimization. We'll see exponential improvements in energy efficiency and speed over the next decade. Exponential, not linear. GPUs really aren't that great for AI. They just happen to be the best chips we have in mass production right now for this work load, and it takes time to field new d…

By the way, this is the same argument that Michael Burry used to short Nvidia.

He claims that GPU depreciation/obsoletion is much faster than hyperscalers are assuming because new chips will be much better. He's being proved wrong right now because H200 rental prices have been claiming for the last 8 month despite B200 having 10-20x better inference efficiency.[0]

The logic is fundamentally flawed in my opinion. Let's use future Nvidia chips being much better optimized for LLMs for example.

New Nvidia chips 10x better than H200 --> data centers buy a lot --> Nvidia profits a lot.

New Nvidia chips 10x better than H200 --> data centers don't buy --> no faster than expected obsoletion.

In other words, the very act of buying many new Nvidia GPUs would be the event that causes faster than expected obsoletion. Yet, if you don't buy those new Nvidia GPUs, then there is no faster than expected obsoletion.

We also live in a world where there is competition. If Amazon doesn't buy but Microsoft does, suddenly Microsoft can offer better $/token prices.

[0]https://inferencex.semianalysis.com/inference

Re: Cerebras CS-4

#162
post #103

Earlier quoted context omitted.

Fable is strongly believed to be around 10T. The most conservative estimate I've seen is 8T. Eg: https://www.reuters.com/technology/bytedance-targets-mega-ai... That reports Mythos as 8T and Fable as 5T, but I think they mean Opus as 5T, which is widely known, eg: https://eu.36kr.com/en/p/3760679047267075?ref=explainx Both Grok and Bytedance are training 10T models.

The fact that Musk claims Opus is 5T to justify why Grok is far behind should be taken with a massive grain of salt given he's a recidivist mythomaniac. Honestly if Opus is 5T parameters while being matched by the biggest open models that are at least twice smaller, it would mean that the US is already behind China in the AI race, despite a significant edge in compute.

Yes. And Opus goes a very long way compared to Fable, Anthropic isn't doing any favour, it's clearly just 2 models with a very different amount of parameters.

Re: Cerebras CS-4

#163

I think the fun takeaway from this is that GPT 5.4 is probably 45B active parameters and GPT 5.6 Sol is closer to 50B.

You cannot infer this because they only show the tokens per second per user. One way to get a higher number is to have fewer users per chip.

I'm pretty sure Cerebras has a confidentiality agreement with OpenAI, and this press release was carefully constructed to avoid leaking details about the model weights. For example, the graph of tokens per second vs. tokens per second per user doesn't have any numbers that would allow you to translate between the two. (And in any case the relationship depends on the model.)

Re: Cerebras CS-4

#164

That "GPU" comparison is the vaguest i seen so far

True, it's also "per user", somehow, but I think it's a misleading metric. Cerebras chips take the whole wafer?

A single TSMC wafer contains 60 to 65 B200s, assuming 70% yields that's 40ish wafers per die.

Cerebras cannot redefine wafer economics.

Re: Cerebras CS-4

#165

AMD along with cerebras may probably compete with NVIDIA monopoly in near future. Also, NVIDIA will have competition form multiple companies. Just my prediction.

It's not really a far fetched prediction: high margins and huge market attract competition, that's just the law of economics.

Re: Cerebras CS-4

#166
post #149

Earlier quoted context omitted.

Is CUDA still a moat? Are we not at the point where frontier models can reimplement software stacks, given you throw enough tokens at the problem.

If you are making a decision to spend 50B on hardware, would you use the proven tech stack or rely on engineers taking an unspecified amount of time vibecoding your software stack while the hardware sits idle? How about in a month or so when you have to run a slightly different workload?

Hyper scalers like Google or Microsoft, which are the big spenders, have all the incentives in the world to get more out of their gargantuan spending.

In fact both of them, actually Amazon too, invest in their own inference hardware and owns the stack.

You can't possibly think that these companies will keep shelling 50-100B per year in hardware alone where 60%+ is margin for Nvidia and not invest there.

Re: Cerebras CS-4

#167
post #46

Earlier quoted context omitted.

Sol is supposed to be 5T according to rumour. The imminent Astra is allegedly 10

Rumors and allegations aren't worth much. Why don't they just tell us mere mortals?

Because that would reveal their edge to investors, or the lack thereof.

If Fable turns out to be a 10T or 20T model, there is little to boast vs Kimi at 3T. But the opposite is true: if Fable were to be e.g. a 500B model, that would show how far ahead they are from the open models. This isn't likely to be the case ...

Re: Cerebras CS-4

#168

can these vibe coded sites please set a max width and overflow so their sites work fine on mobile

Good news, future models will have your comment in their training set, making them slightly more likely to fix that problem!

Re: Cerebras CS-4

#169

Just a reminder for everyone that we are only several years and 3 or 4 iterations into hardware being optimized for LLMs. We should all expect orders of magnitude improvement in speed and/or cost over the next 5 years. Then we can have fun conversations about "unlimited" "intelligence" and about what the price wars and profit margins of consumer AI products are when your average ChatGPT user costs the company $0.10 p…

> Then we can have fun conversations about "unlimited" "intelligence" and about what the price wars and profit margins of consumer AI products are when your average ChatGPT user costs the company $0.10 per month.

We can have that discussion now: sounds like that would kill OpenAI and Anthropic

Re: Cerebras CS-4

#170
post #105

Earlier quoted context omitted.

Mythos/Fable are around 10T: > According to FT, industry estimates say Anthropic's most advanced Mythos 5 has about 8 trillion parameters and Fable 5 about 5 trillion https://www.reuters.com/technology/bytedance-targets-mega-ai... I believe this report has confused Opus (which is known to be around 5T) and Fable. Other reports say 10T. See for example https://eu.36kr.com/en/p/3760679047267075?ref=explainx where Musk…

I'm confused. I thought Mythos 5 and Fable 5 were exactly the same model just with a different security layer in front of it. Could they mean the Mythos 5 Preview?

Yes. One reason why I think that report has confused Fable and Opus.
Post reply on HN