My main question is whether when put into practical use, this can be measured in tokens/second, or more like 1 token per minute... I have seen locally hosted LLM that are as slow as 1 tok/second still be very useful if you give it a project to do something overnight and metaphorically walk away from it, check back with what it has done in 6 or 8 hours. 0.05 to 0.1 tok/s on the other hand, as reported in the URL for t…
> on hardware that ordinary people can afford These days, can "ordinary people" afford 24GB of ram and half a TB of NVME ssd? sigh
You can, right now, buy a brand new Mini-PC at or above this spec for $600 at retail [1]
Of course, if you want it in a desktop format with a much faster CPU, its going to cost you more.
[1]: https://www.amazon.com/GMKtec-M6-Ultra-Upgraded-Computers/dp...