Earlier quoted context omitted.
The implementation absolutely can influence the outputs. If you have a sloppy implementations which somehow accumulates a lot of error in it's floating point math, you will get worse results. It's rarely talked about, but it's a real thing. Floating point addition and multiplication is non-associative and the order of operations affects the correctness and performance. Developers might (unknowningly) trade performanc…
I thought all current implementations accumulate into a fp32 instead of accumulating in fp16.
Does anyone have experience with higher-precision matmul and whether it is worthwhile?