Also does this lower total cost depend on SRAM being available for DRAM prices?
What makes SRAM so much more expensive than DRAM?
31–40 of 91 posts
Also does this lower total cost depend on SRAM being available for DRAM prices?
What makes SRAM so much more expensive than DRAM?
How much have they optimized the software here? Is it tinygrad level optimization? Also does this lower total cost depend on SRAM being available for DRAM prices? What makes SRAM so much more expensive than DRAM?
How much have they optimized the software here? Is it tinygrad level optimization? Also does this lower total cost depend on SRAM being available for DRAM prices? What makes SRAM so much more expensive than DRAM?
How much have they optimized the software here? Is it tinygrad level optimization? Also does this lower total cost depend on SRAM being available for DRAM prices? What makes SRAM so much more expensive than DRAM?
Earlier quoted context omitted.
Large language models would need tens or hundreds of gigabytes of SRAM. Pretty sure the enormous cost for this makes the approach economically unfeasible.
Cerebras Wafer Scale Engine has 40GB of onboard SRAM using TSMC 7nm. It uses the entire wafer as the chip. Costs millions per chip though. Source: https://www.anandtech.com/show/16626/cerebras-unveils-wafer-...
Earlier quoted context omitted.
Cerebras Wafer Scale Engine has 40GB of onboard SRAM using TSMC 7nm. It uses the entire wafer as the chip. Costs millions per chip though. Source: https://www.anandtech.com/show/16626/cerebras-unveils-wafer-...
Wafer scale integration is a dead end technology. The engineering issues are just too great.
Earlier quoted context omitted.
Wafer scale integration is a dead end technology. The engineering issues are just too great.
Do you have a source?
Look at how successful AMD's chiplet strategy has been. Chiplets sidestep the yield problems. Wafer scale amplifies them hundred or thousand fold.
Nothing in the industry is designed to work with wafer scale products, so everything has to be custom made. Yes this is a chicken-egg problem, but it's going to be expensive to get any sort of momentum. The silicon industry is extremely conservative.
It's sexy and enticing. If someone can make it work that's awesome. I will remain skeptical though.
This seems like a pretty bad paper. Their headline claim that they are 300x faster than an A100 at serving GPT-3 uses obviously wrong numbers for how fast A100s can run GPT-3. They seem to have misread the DeepSpeed Inference paper and claim that the best throughput on GPT-3 sized models was 18 tok/s, but if you look at figure 8 on page 11 of the paper [1], it shows that they are able to achieve ~74 teraflops on serv…
Earlier quoted context omitted.
Large language models would need tens or hundreds of gigabytes of SRAM. Pretty sure the enormous cost for this makes the approach economically unfeasible.
Cerebras Wafer Scale Engine has 40GB of onboard SRAM using TSMC 7nm. It uses the entire wafer as the chip. Costs millions per chip though. Source: https://www.anandtech.com/show/16626/cerebras-unveils-wafer-...