In case anyone is wondering, yes, there is a cost when a model is quantized. https://oobabooga.github.io/blog/posts/perplexities/ Essentially, you lose some accuracy and there might be some weird answers and probably more likely to go off the rail and hallucinate. But the quality loss is lower the more parameters you have. So for very large model sizes the differences might be negligible. Also, this is the cost of in…
Sadly it didn't. It talked about "perplexities" and showed some floating point numbers.
I want to see examples like "here's a prompt against a model and the same prompt against a quantized version of that model, see how they differ."