Cool. Now run TerminalHard and compare to unquantized 27B. KLD of 1%, or similar error metric that multiplies, on 10000 tokens would give accumulated error of 2,000,000%
I don't think you can extrapolate that measurement across multiple sequential draws like that. We presumably are comparing against a single trajectory rather than a tree of trajectories. So once we make the wrong choice and step off of the blessed path, we have no way to assign a ranking to the next token; it's error is undefined. I've seen LLMs self correct in chains of thought ("because of foo and bar, I need to...…
but wait, the models constantly go back and forth on these things in their thinking traces, so it is unclear which self correcting is actually correct