Does it save all conversations and let me revisit them later? I use MLC Chat to run Mistral 7B on my iPhone at the moment, but the lack of conversation history is a real nuisance: https://apps.apple.com/us/app/mlc-chat/id6448482937
In your experience, how could these local LLMs become snappier than using streamed API calls? How far are they if not? How soon do you guess they’ll get there? I understand the motivation includes factors other than performance, I’m just curious about performance as it applies to UX.
There is no expectation that phones will ever be comparable in performance for LLMs.
Mistral runs at a decent clip on phones, but we’re talking like 11 tokens per second, not hundreds of tokens per second.
Server-based models tend to be only slightly faster than Mistral on my phone because they’re usually running much larger, much more accurate/useful models. Models which currently can’t fit onto phones.
Running models locally is not motivated by performance, except if you’re in places without reliable internet.