Better sell all nvidia stocks. Once these chips are common there is no need anymore for GPUs in training super large AI models.
I trust that gamers will outlast every hype, be it crypto or AI.
41–50 of 93 posts
Better sell all nvidia stocks. Once these chips are common there is no need anymore for GPUs in training super large AI models.
I trust that gamers will outlast every hype, be it crypto or AI.
Better sell all nvidia stocks. Once these chips are common there is no need anymore for GPUs in training super large AI models.
But can it run doom?
- Interconnect between WSE-2's chips in the cluster was 150GB/s, much lower than NVIDIA's 900GB/s. - non-sparse fp16 in WSE-2 was 7.5 tflops (about 8 H100s, 10x worse performance per dollar) Does anyone know the WSE-3 numbers? Datasheet seems lacking loads of details Also, 2.5 million USD for 1 x WSE-3, why just 44GB tho???
You can order one with 1.2 Petabytes of external memory. Is that enough?
"External memory: 1.5TB, 12TB, or 1.2PB"
https://www.cerebras.net/press-release/cerebras-announces-th...
"214Pb/s Interconnect Bandwidth"
Is there a reason it's not roughly a disc if they use the whole wafer ? They could have 50% more surface.
According to the company, the new chip will enable training of AI models with up to 24 trillion parameters. Let me repeat that, in case you're as excited as I am: 24. Trillion. Parameters. For comparison, the largest AI models currently in use have around 0.5 trillion parameters, around 48x times smaller. Each parameter is a connection between artificial neurons . For example, inside an AI model, a linear layer that…
So only 4-20 of these systems are necessary to match the human brain. No?
Reposting the CS-2 teardown in case anyone missed it. The thermal and electrical engineering is absolutely nuts: https://vimeo.com/853557623 https://web.archive.org/web/20230812020202/https://www.youtu... (Vimeo/Archive because the original video was taken down from YouTube)
Reposting the CS-2 teardown in case anyone missed it. The thermal and electrical engineering is absolutely nuts: https://vimeo.com/853557623 https://web.archive.org/web/20230812020202/https://www.youtu... (Vimeo/Archive because the original video was taken down from YouTube)
20,000 amps 200,000 electrical contacts 850,000 cores and that's the "old" one. wow.
edit: thanks people, makes sense now!
If you were to add up all transistors fabricated worldwide, up until , such that total roughly matches the # on this beast, what year would you arrive? Hell, throw in discrete transistors if you want. How many early supercomputers / workstations etc would that include? How much progress did humanity make using all those early machines (or any transistorized device!) combined?
4004 from the 1970s used 2300 transistors, so it would have needed to sell billions.
Pentium from 1990s had 3M transistors, so it could hit our target by selling a million units.
I'm betting (without much research) that the Pentium line alone sold millions, and the industry as a whole could hit those numbers about 5 years earlier.
- Interconnect between WSE-2's chips in the cluster was 150GB/s, much lower than NVIDIA's 900GB/s. - non-sparse fp16 in WSE-2 was 7.5 tflops (about 8 H100s, 10x worse performance per dollar) Does anyone know the WSE-3 numbers? Datasheet seems lacking loads of details Also, 2.5 million USD for 1 x WSE-3, why just 44GB tho???
>> why just 44GB tho??? You can order one with 1.2 Petabytes of external memory. Is that enough? "External memory: 1.5TB, 12TB, or 1.2PB" https://www.cerebras.net/press-release/cerebras-announces-th... "214Pb/s Interconnect Bandwidth" https://www.cerebras.net/product-system/