They say there are no stupid questions, so here is mine: If there are Billions of parameters in the SOTA models, how do we argue that they are not over fitting?
That's section 4 of OP.
Clearly, the authors have given careful considerations to the issue of contamination and have provided reasonable analysis and a careful argument regarding over fitting the existing benchmarks.
On the other I was wondering if the authors would like to consider purposefully creating a type of "out of sample data" for "creative evaluation"? Of course, GPT is no stranger to creativity, so it would be a fascinating challenge to come up with methods to create such datasets that are truly creative and challenge GPT-{N} to prove its mettle.
For example, would it be possible to engage a really good creative writer* along with a highly experienced school teacher to take on the Reading Comprehension task and create few "tricky" evaluation samples that not only go above and beyond the contamination objections but also challenge the human intelligence to be careful not to fall into common traps?
This way lies a different evaluation metric - a subjective one perhaps, but it's a start. Just a thought experiment - that's all.
* so that they can come up with new ways to trick GPT/humans a teacher knows the common mistakes the average student makes
Edit: Duh, my head immediately screamed GANs the moment I pressed submit, lol. But I am not sure if GANs make sense for NLP tasks. Like do they make sense if humans/domain experts try to solve them?