If Apple would wake up to what's happening with llama.cpp etc then I don't see such a market in paying for remote access to big models via API, though it's currently the only game in town. Currently a Macbook has a Neural Engine that is sitting idle 99% of the time and only suitable for running limited models (poorly documented, opaque rules about what ops can be accelerated, a black box compiler [1] and an apparent…
If I were Apple I'd be thinking about the following issues with that strategy: 1. That RAM isn't empty, it's being used by apps and the OS. Fill up 64GB of RAM with an LLM and there's nothing left for anything else. 2. 64GB probably isn't enough for competitive LLMs anyway. 3. Inferencing is extremely energy intensive, but the MacBook / Apple Silicon brand is partly about long battery life. 4. Weights are expensive t…
2. see above
Should be cheap, or why else are Samsung, Micron and Kioxia whining about losses?
Maybe go for something like Optane memory while doing so.