Earlier quoted context omitted.
Compute Express Link (CXL) should mostly solve limited RAM with CPU: 1) Compute Express Link (CXL): https://en.wikipedia.org/wiki/Compute_Express_Link PCIe vs. CXL for Memory and Storage: https://news.ycombinator.com/item?id=38125885
Gigabytes per second? What is this, bandwidth for ants? My years old pleb tier non-HBM GPU has more than 4 times the bandwidth you would get from a PCIe Gen 7 x16 link, which doesn't even officially exist yet.
[1] Forget ChatGPT: why researchers now run small AIs on their laptops:
https://news.ycombinator.com/item?id=41609393
[2] Welcome to LLMflation – LLM inference cost is going down fast:
https://a16z.com/llmflation-llm-inference-cost/
[3] New LLM optimization technique slashes memory costs up to 75%: