Earlier quoted context omitted.
Tuning the model output to perform better on certain prompts is not the same as improving the model. It's valid to worry that the model makers are gaming the benchmarks. If you think that's happening and you want to personally figure out which models are really the best, keeping some prompts to yourself is a great way to do that.
Who’s going out of their way to optimize for random HNers informal benchmarks?
So I would guess every single AI being made currently