Earlier quoted context omitted.
I'm no LLM evangelist, far from it, but I expect models of similar quality to the current bleeding-edge, will be freely runnable on consumer hardware within 3 years. Future bleeding-edge models may well be more expensive than current ones, who knows.
How do the best models that can run on say a single 4090 today compare to GPT 3.5?
https://llm-stats.com/models/compare/gpt-3.5-turbo-0125-vs-q...