Chiplet ASIC supercomputers for LLMs like GPT-4
1–10 of 91 posts
Re: Chiplet ASIC supercomputers for LLMs like GPT-4
#2Re: Chiplet ASIC supercomputers for LLMs like GPT-4
#3This development presents a more compelling case that we are in fact on the precipice of larger LLMs being able to serve everyone for cheap. Still not really convinced by the AGI argument, but this does spook me. Overall though very cool.
Re: Chiplet ASIC supercomputers for LLMs like GPT-4
#4Re: Chiplet ASIC supercomputers for LLMs like GPT-4
#5Re: Chiplet ASIC supercomputers for LLMs like GPT-4
#6They claim the investment will be justified for a 1.5 year life span of the system. But LLMs are changing and improving at a much faster speed that 1.5 years feels like centuries!
Re: Chiplet ASIC supercomputers for LLMs like GPT-4
#7A key architectural feature to achieve this is the ability to fit all model parameters inside the on-chip SRAMs of the chiplets to eliminate bandwidth limitations. Doing so is non-trivial as the amount of memory required is very large and growing for modern LLMs
...
On-chip memories such as SRAM have better read latency and read/write energy than external memories such as DDR or HBM but require more silicon per bit. We show this design choice wins in the competition of TCO per performance for serving large generative language models but requires careful consideration with respect to the chiplet die size, chiplet memory capacity and total number of chiplets to balance the fabrication cost and model performance (Sec.3.2.2) We observe that the inter-chiplet communication issues can be effectively mitigated through proper software-hardware co- design leveraging mapping strategies such as tensor and pipeline model parallelism
Re: Chiplet ASIC supercomputers for LLMs like GPT-4
#8Re: Chiplet ASIC supercomputers for LLMs like GPT-4
#9They claim the investment will be justified for a 1.5 year life span of the system. But LLMs are changing and improving at a much faster speed that 1.5 years feels like centuries!
"Moving fast" may take on a whole new meaning and I'd put money on the rate of iteration soon being beyond the vast majority's comprehension (myself included).
Re: Chiplet ASIC supercomputers for LLMs like GPT-4
#1094x cost improvement over GPU and 15x TPU is insane, but fits right in there with performance gains seen in Moore's Law. This development presents a more compelling case that we are in fact on the precipice of larger LLMs being able to serve everyone for cheap. Still not really convinced by the AGI argument, but this does spook me. Overall though very cool.