Earlier quoted context omitted.
A major reason I haven’t really tried any of these things (despite thinking they are vaguely neat). I think I will wait until… 2026, second half, most likely. At least I’ll check if we have local models and hardware that can run them nicely, by then. Hats off to the folks who have decided to deal with the nascent versions though.
It is completely unreasonable to buy the hardware to run a local model and only use it 1% of the time. It will be unreasonable in 2026 and probably very long after that. Maybe s.th. like a collective that buys the gpu's together and then uses them without leaking data can work.
maybe 128gb of vram becomes the new mid tier model and most llms can fit into this nicely and do everything one wants in an llm
given how fast llms are progressing it wouldn’t surprise me if we reach this point by 2030