I wonder how long it would take a "normal" coding prompt to go thru a "1T-2T" models on a medium performance consumer desktop. Hours? Days?
https://www.youtube.com/watch?v=9TyJ9s26ylc acoording to this video, around 30t/s, if in 2027 some 256gb ai box arrive, the performance will be even better
I did watch your vid: it is for quantized deepseek models. I would look for full weight inference for maximum quality.
For instance, do we have similar benchmarks on the latest frontier open weight KIMI model at several tera-weights (I would be interested at coding prompts)?