Earlier quoted context omitted.
1. You're working backwards from a desire to buy more RAM to try and find uses for it. I'm really not I had no desire at all until a couple of weeks ago. Even now not so much since it wouldn't be very useful to me But the current LLM business model where there are a small number of API providers, and anything built using this new tech is forced into a subscription model... I don't see it sustainable, and I think the…
I think llama.cpp will die soon because the only models you can run with it are derivatives of a model that Facebook never intended to be publicly released, which means all serious usage of it is in a legal limbo at best and just illegal at worst. Even if you get a model that's clean and donated to the world, the quality is still not going to be competitive with the hosted models. And yes I've played with it. It was/…
(B) llama.cpp supports gpt4all, which states that its working on fixing your concern. This is from their README:
Roadmap Short Term
- Train a GPT4All model based on GPTJ to alleviate llama distribution issues.