Earlier quoted context omitted.
It's rumored fable is around that 10T number
If this is true, it's even more impressive that some of the open weight models that are <3.5T in size, approx 33% of its size, are within a few points of it in the artificial analysis leaderboard.
Cerebras CS-4
81–90 of 281 posts
Re: Cerebras CS-4
#82Re: Cerebras CS-4
#83Re: Cerebras CS-4
#84Re: Cerebras CS-4
#85Earlier quoted context omitted.
Information about RAM type/size and connection topology of the RAM to be used for context cache seems to be conspicuously absent from the slick looking marketing materials.
There's a few more details at the bottom of this page: https://www.cerebras.ai/blog/introducing-cerebras-cs-4 44GB on-chip-sram * 3 chips. Per chip: 43.2 PB/s memory access + 53.5 PB/s on-chip fabric bandwidth + 2.4 Tbits/s "IO" bandwidth (I think that means their RoCE v2 RDMA over Ethernet interface). I suspect there might be a certain amount of customization for how much RAM they attach when you order it.
This is 1/3rd blackwells nvlink c2c bandwidth already. Not too bad. We can make KV cache offload work with that I suppose.
If magically KV cache was not an issue, pipeline parallelism on cerebras can be quite pleasant. As for the KV cache offload, I have hopes their CPO solution they're trying with that canadian company ends up bearing fruit.
Re: Cerebras CS-4
#86Re: Cerebras CS-4
#87Five years from now, I don't know why anyone will still be using Nvidia for inference. Note that Cerebras is for inference only, not for training. I understand that Cerebras has competition, but this bodes even more poorly for Nvidia for inference. Nvidia may still have a role to play for training, however.
Re: Cerebras CS-4
#88Just a reminder for everyone that we are only several years and 3 or 4 iterations into hardware being optimized for LLMs. We should all expect orders of magnitude improvement in speed and/or cost over the next 5 years. Then we can have fun conversations about "unlimited" "intelligence" and about what the price wars and profit margins of consumer AI products are when your average ChatGPT user costs the company $0.10 p…
Hence why taalas was one of the best strategic acquisitions of the year. I'm honestly baffled they were not acquired by somebody else (sorry AMD).
The absolute worst market time to etch a model to a chip is right now (very rapid iteration). There is no scenario where they can keep up. The Taalas approach will be viewed as comically foolish within just a few years.
Cerebras will win in terms of approach.
It's 1998: hey, I can drastically speed up your web service, let's etch it right to silicon.
Re: Cerebras CS-4
#89Earlier quoted context omitted.
TPUs and ASICs run in data centers too. Your argument only holds true if there's some satisfied limit to demand for inference. If not, data centers will continue to spring up to host more and more agents. Even if agents were running on hardware and software as efficient as the human brain, its conceivable we want trillions of them running at any given time which would require data center scale.
Everything has some satisfied limit to demand, often depending on the price. If you assume there will never be any satisfied limit to demand for inference at any price you can justify any investment.
Unlimited.
What has been the limit to electricity demand globally?
Unlimited.
We can't get enough and never will. Costs have to become pretty severe to turn back the demand as well.