Earlier quoted context omitted.
> on a decent speed But you said 7-9 tokens/second , that's not a decent speed. I'm not an expert by all means but in my local experiments, less than 12 to 16 tps is too slow.
I think that any workflow that requires the user to stare at the tokens being generated live is using it wrong. Delegate, don't stare! https://mikeveerman.github.io/tokenspeed/?rate=10&mode=text You think of an idea that you want to have the LLM process, queue it up, and go back to what you were doing. Once you've finished reading the next article on HN about a 5 tps Xeon, your task will be complete. It's kind of lik…
This is how I used to think about my 3D printer, but FWIW the way my actual thinking and planning works, print speed really matters. Not for the final print, but for iterative work and test parts, it is obvious that either having a fast printer helps. Having multiple slow printers also helps, but there are only so many areas of a design you can iterate on at once.
At the moment my own LLM use is experimental and iterative, and I definitely favour the faster MoE models for much of what I am doing, even if I might in principle prefer to get the final work done in the slower ones.