> Due to cost limitations we had to limit crafty to 2 seconds of analysis time per move A grandmaster with standard time controls could defeat a 2-second limited Crafty. So how do you know you're finding true blunders, and not simply positions that the engine evaluates incorrectly?
This is definitely the biggest limitation of our approach right now and there are certainly some things that we counted as blunders that aren't true blunders. We're working on rectifying this by doing another pass with a better engine and more time to analyze. That said we tested this on a smaller set of games by comparing it to results from better engines and found that only a very small number of moves tricked craf…
The results of the cross-validation you mentioned would be interesting as well.