Earlier quoted context omitted.
Wait, I may be missing something here. These benchmarks are gathered by having models play each other, and the second illegal move forfeits the game. This seems like a flawed method as the models who are more prone to illegal moves are going to bump the ratings of the models who are less likely. Additionally, how do we know the model isn’t benchmaxxed to eliminate illegal moves. For example, here is the list of games…
That’s a devastating benchmark design flaw. Sick of these bullshit benchmarks designed solely to hype AI. AI boosters turn around and use them as ammo, despite not understanding them.
https://arxiv.org/abs/2403.15498