Earlier quoted context omitted.
The difference is small, UNTIL you get to 4 bit quantization, where the model is noticeably dumber. 8 bits, imo, is the minimum.
Some parameters would be more sensitive than others I suppose? So could you use 4 bits for most, and 8 bits, or even 16, for the remaining?
I do wonder if it would be possible to have the model determine during training how important each parameter is, while maybe rewarding it for having more small parameters?