Live data from Hacker News

Lexar Wants to Offload Local AI Models to SSD Amid the RAMpocalypse

techpowerup.com

1–3 of 3 posts

Re: Lexar Wants to Offload Local AI Models to SSD Amid the RAMpocalypse

#2
There have also been proposals to use flash memory in inference accelerators instead of DRAM. You can make high bandwidth flash using the same stacking technique used for HBM DRAM.

It is obviously unsuitable for training because of limited write cycles. But the read bandwidth is decent, and the density/$ is much better.