One thing I don't understand about their architecture is that they have spent so much effort building this monster of a chip, but if you are going to do something crazy, why not work on processing in memory instead? At least for transformers you will primarily be bottlenecked on matrix multiplication and almost nothing else, so you only need to add a simple matrix vector unit behind your address decoder and then almo…
4T transistors, one giant chip (Cerebras WSE-3) [video]
61–70 of 93 posts
Re: 4T transistors, one giant chip (Cerebras WSE-3) [video]
#62Reposting the CS-2 teardown in case anyone missed it. The thermal and electrical engineering is absolutely nuts: https://vimeo.com/853557623 https://web.archive.org/web/20230812020202/https://www.youtu... (Vimeo/Archive because the original video was taken down from YouTube)
20,000 amps 200,000 electrical contacts 850,000 cores and that's the "old" one. wow.
The peak power average of the 7950X3D is roughly 150W[3], which means if you could somehow run all 450 CPUs (900 dies) at peak, they'd consume around 68kW.
edit: I forgot about the IO die which contains the memory controller, so that will suck some power as well. So if we say 50W for that and 50W for the CPU dies, that's 45kW.
That's assuming you get a "clean wafer" with all dies working, not "just" the 80% yield or so.
[1]: https://www.techpowerup.com/cpu-specs/ryzen-9-7950x3d.c3024
[2]: https://www.anandtech.com/show/15219/early-tsmc-5nm-test-chi...
[3]: https://www.tomshardware.com/reviews/amd-ryzen-9-7950x3d-cpu...
Re: 4T transistors, one giant chip (Cerebras WSE-3) [video]
#63Not trying to sound critical, but is there a reason to use 4B,000 vs 4T?
Quote:
Billion is a word for a large number, and it has two distinct definitions:
1,000,000,000, i.e. one thousand million, or 10^9 (ten to the ninth power), as defined on the short scale. This is now the most common sense of the word in all varieties of English; it has long been established in American English and has since become common in Britain and other English-speaking countries as well.
1,000,000,000,000, i.e. one million million, or 10^12 (ten to the twelfth power), as defined on the long scale. This number is the historical sense of the word and remains the established sense of the word in other European languages. Though displaced by the short scale definition relatively early in US English, it remained the most common sense of the word in Britain until the 1950s and still remains in occasional use there.
Re: 4T transistors, one giant chip (Cerebras WSE-3) [video]
#64Better sell all nvidia stocks. Once these chips are common there is no need anymore for GPUs in training super large AI models.
This chip does not outperform NVIDIA on key metrics. Economics of scale are unfavorable. Software is exotic. I trust that gamers will outlast every hype, be it crypto or AI.
Re: 4T transistors, one giant chip (Cerebras WSE-3) [video]
#65- Interconnect between WSE-2's chips in the cluster was 150GB/s, much lower than NVIDIA's 900GB/s. - non-sparse fp16 in WSE-2 was 7.5 tflops (about 8 H100s, 10x worse performance per dollar) Does anyone know the WSE-3 numbers? Datasheet seems lacking loads of details Also, 2.5 million USD for 1 x WSE-3, why just 44GB tho???
44GB is the SRAM on a single device, comparable to the 50MB of L2 on the H100. There is also a lot of directly attached DRAM.
Re: 4T transistors, one giant chip (Cerebras WSE-3) [video]
#66Better sell all nvidia stocks. Once these chips are common there is no need anymore for GPUs in training super large AI models.
I would be more worried about the fact that next year every CPU is going to ship with some kind of AI accelerator already integrated to the die, which means the only competitive differentiation boils down to how much SRAM and memory bandwidth your AI accelerator is going to have. TOPS or FLOPS will become an irrelevant differentiator.
Re: 4T transistors, one giant chip (Cerebras WSE-3) [video]
#67- Interconnect between WSE-2's chips in the cluster was 150GB/s, much lower than NVIDIA's 900GB/s. - non-sparse fp16 in WSE-2 was 7.5 tflops (about 8 H100s, 10x worse performance per dollar) Does anyone know the WSE-3 numbers? Datasheet seems lacking loads of details Also, 2.5 million USD for 1 x WSE-3, why just 44GB tho???
Re: 4T transistors, one giant chip (Cerebras WSE-3) [video]
#68But can it run doom?
Re: 4T transistors, one giant chip (Cerebras WSE-3) [video]
#69Is there a reason it's not roughly a disc if they use the whole wafer ? They could have 50% more surface.