Live data from Hacker News

Cerebras CS-4

cerebras.ai

141–150 of 281 posts

Re: Cerebras CS-4

#142
post #52

It would be even better if a version available to individual users were released soon.

I'd like to see a consumer version too, I don't need a whole rack of them. I probably can't even afford one gpu-sized one

Re: Cerebras CS-4

#143

AMD along with cerebras may probably compete with NVIDIA monopoly in near future. Also, NVIDIA will have competition form multiple companies. Just my prediction.

Like NVIDIA bought Groq, AMD might do well buying Cerebras.

They are Ex AMD employees. Nvidia tried back, but they rejected.

Re: Cerebras CS-4

#144
post #134

Earlier quoted context omitted.

Taalas will be one of the great disaster investments of the early AI era. It'll be a near total write-down. The absolute worst market time to etch a model to a chip is right now (very rapid iteration). There is no scenario where they can keep up. The Taalas approach will be viewed as comically foolish within just a few years. Cerebras will win in terms of approach. It's 1998: hey, I can drastically speed up your web…

I would still pay ~500 for a chip that runs 10kt/s of a ~100b model on my machine even if the half life is 6months. I m sure my employer would too. Qwwen3.5 122b was released 6 months ago and is still best in class overall 100-140 B param model.

10000%

Re: Cerebras CS-4

#145
post #3

Just a reminder for everyone that we are only several years and 3 or 4 iterations into hardware being optimized for LLMs. We should all expect orders of magnitude improvement in speed and/or cost over the next 5 years. Then we can have fun conversations about "unlimited" "intelligence" and about what the price wars and profit margins of consumer AI products are when your average ChatGPT user costs the company $0.10 p…

This is part of why I think the data center build-out is a bubble. We've barely scratched the surface when it comes to hardware optimization. We'll see exponential improvements in energy efficiency and speed over the next decade. Exponential, not linear. GPUs really aren't that great for AI. They just happen to be the best chips we have in mass production right now for this work load, and it takes time to field new d…

Exactly right, and nVidia is protecting their moat through business practices rather than genuine product innovation.

Re: Cerebras CS-4

#146
post #105

Earlier quoted context omitted.

Mythos/Fable are around 10T: > According to FT, industry estimates say Anthropic's most advanced Mythos 5 has about 8 trillion parameters and Fable 5 about 5 trillion https://www.reuters.com/technology/bytedance-targets-mega-ai... I believe this report has confused Opus (which is known to be around 5T) and Fable. Other reports say 10T. See for example https://eu.36kr.com/en/p/3760679047267075?ref=explainx where Musk…

> I believe this report has confused Opus (which is known to be around 5T) and Fable. 5T for Opus feels quite high though. DeepSeek V4 Pro is a mere 1.6T and often described as a match with Opus in overall quality. Even the largest open models in common use are around 2.8T.

> and often described as a match with Opus in overall quality

It's not. Idk about who has more T's but, unfortunately, DS4 pro is not a match to Opus, at least not Opus 4.8.

Re: Cerebras CS-4

#147
post #40

Earlier quoted context omitted.

It's rumored fable is around that 10T number

Fable is most definitely nowhere near 10T. The cost to train and infer that would be insane, even by today's standards.

> The cost to train and infer that would be insane, even by today's standards.

This assumption is likely what has led to the erroneous failure.

Enterprise compute per rack has scaled multiple fold in the last 3-5 years. Alongside the training efficiency gains & datacenter scale increases, even 50T+ is well within reach at the top end.

Re: Cerebras CS-4

#149
post #80

Earlier quoted context omitted.

Maybe, but GPU is just one aspect of NVIDIA's dominance. If you are buying Vera Rubin GPUs, you're getting an NVL72 rack, which is only one of several racks that you're probably buying. You'll also need your NVIDIA racks with NVIDIA networking & storage gear, too. At the end of the day, they're "vertically integrated" for your accelerated computing data center (e.g. the "AI Factory"). This doesn't even count the soft…

Is CUDA still a moat? Are we not at the point where frontier models can reimplement software stacks, given you throw enough tokens at the problem.

If you are making a decision to spend 50B on hardware, would you use the proven tech stack or rely on engineers taking an unspecified amount of time vibecoding your software stack while the hardware sits idle?

How about in a month or so when you have to run a slightly different workload?

Re: Cerebras CS-4

#150
post #113

Earlier quoted context omitted.

What about general knowledge you can get out of it before hallucinations start?

I do not rely on any LLM of any size for general knowledge baked into the weights, they all hallucinate and that is the wrong way to hold them imo I think there is some merit in that smaller models cannot memorize so much of the training data, i.e. that they are less likely to do copyright infringement, and by analogy not having memorized SDK / API surfaces that have since changed from the training data

> I do not rely on any LLM of any size for general knowledge baked into the weights

You have to rely on it to a certain level for agentic/coding work, presuming that's the general subject we're talking about here... For instance I recently encountered a project where it would have been a lot worse if the LLM didn't already know "what is" xterm.js and a bunch of its associated npm-related/node related software. If it was still smart but had to google and find results for everything it would have been a lot more time consuming and risked sending it down a wrong path.

Post reply on HN