>For benchmarks, there should be held out tests that don't get distributed to (say) compiler authors until after the results are published.
The way you really "solve" the problem is to have a third-party test organization run the benchmark and only provide the final result--and you only get one shot for a given generation of hardware/software. (This is sort of how standardized tests work.)
The problem is that you're now benchmarking with a totally opaque test that you have to trust is reasonably representative of the type of workload it claims to represent as well as trust that the benchmark isn't favoring specific vendors or design decisions that are irrelevant to real world workloads.
I assume there have been benchmarks at various times that were similar to this. Certainly, I'm pretty sure source code for benchmark suites hasn't always been available.