Live data from Hacker News

Cerebras CS-4

cerebras.ai

81–90 of 281 posts

Re: Cerebras CS-4

#81
post #40

Earlier quoted context omitted.

It's rumored fable is around that 10T number

If this is true, it's even more impressive that some of the open weight models that are <3.5T in size, approx 33% of its size, are within a few points of it in the artificial analysis leaderboard.

You want to take a look at the "Scaling Laws" paper, so you can extrapolate from these numbers.

Re: Cerebras CS-4

#83
Cerebras should slowly also move to dgx/ryzen market for a desktop version for masses at affordable price yet providing substantial tokens/second on desktop

Re: Cerebras CS-4

#84
The comparison seems incomplete. CS‑4 is a full rack-scale system with three wafer-scale processors, but the exact GPU models, GPU count, power consumption, price information are not disclosed. We still don't know if buying a multi-GPU rack (or racks) is cheaper and/or more efficient in power. The fact that they didn't disclose these numbers makes me believe that the numbers are not in their favor. And personally, makes me see them as disingenuous.

Re: Cerebras CS-4

#85
post #50

Earlier quoted context omitted.

Information about RAM type/size and connection topology of the RAM to be used for context cache seems to be conspicuously absent from the slick looking marketing materials.

There's a few more details at the bottom of this page: https://www.cerebras.ai/blog/introducing-cerebras-cs-4 44GB on-chip-sram * 3 chips. Per chip: 43.2 PB/s memory access + 53.5 PB/s on-chip fabric bandwidth + 2.4 Tbits/s "IO" bandwidth (I think that means their RoCE v2 RDMA over Ethernet interface). I suspect there might be a certain amount of customization for how much RAM they attach when you order it.

They have managed to make the external link 300GBps/2us. Cs3 was 150/5.

This is 1/3rd blackwells nvlink c2c bandwidth already. Not too bad. We can make KV cache offload work with that I suppose.

If magically KV cache was not an issue, pipeline parallelism on cerebras can be quite pleasant. As for the KV cache offload, I have hopes their CPO solution they're trying with that canadian company ends up bearing fruit.

Re: Cerebras CS-4

#86
post #46

Earlier quoted context omitted.

Sol is supposed to be 5T according to rumour. The imminent Astra is allegedly 10

Rumors and allegations aren't worth much. Why don't they just tell us mere mortals?

Why would they? What the upside, for them?

Re: Cerebras CS-4

#87

Five years from now, I don't know why anyone will still be using Nvidia for inference. Note that Cerebras is for inference only, not for training. I understand that Cerebras has competition, but this bodes even more poorly for Nvidia for inference. Nvidia may still have a role to play for training, however.

Nvidia is at this time a pretty well run company tech wise. They are going to keep iterating on the inferencing hardware stack over the next five years too.

Re: Cerebras CS-4

#88

Just a reminder for everyone that we are only several years and 3 or 4 iterations into hardware being optimized for LLMs. We should all expect orders of magnitude improvement in speed and/or cost over the next 5 years. Then we can have fun conversations about "unlimited" "intelligence" and about what the price wars and profit margins of consumer AI products are when your average ChatGPT user costs the company $0.10 p…

Hence why taalas was one of the best strategic acquisitions of the year. I'm honestly baffled they were not acquired by somebody else (sorry AMD).

Taalas will be one of the great disaster investments of the early AI era. It'll be a near total write-down.

The absolute worst market time to etch a model to a chip is right now (very rapid iteration). There is no scenario where they can keep up. The Taalas approach will be viewed as comically foolish within just a few years.

Cerebras will win in terms of approach.

It's 1998: hey, I can drastically speed up your web service, let's etch it right to silicon.

Re: Cerebras CS-4

#89

Earlier quoted context omitted.

TPUs and ASICs run in data centers too. Your argument only holds true if there's some satisfied limit to demand for inference. If not, data centers will continue to spring up to host more and more agents. Even if agents were running on hardware and software as efficient as the human brain, its conceivable we want trillions of them running at any given time which would require data center scale.

Everything has some satisfied limit to demand, often depending on the price. If you assume there will never be any satisfied limit to demand for inference at any price you can justify any investment.

Looking back nearly 80 years, what has been the limit to transistor demand so far?

Unlimited.

What has been the limit to electricity demand globally?

Unlimited.

We can't get enough and never will. Costs have to become pretty severe to turn back the demand as well.

Post reply on HN