Interesting that in the video, there is an admission that they have been targeting this benchmark. A comment that was quickly shut down by Sam. A bit puzzling to me. Why does it matter ?
It matters to extent that they want to market this as general intelligence, not as a collection of narrow intelligences (math, competitive programming, ARC puzzles, etc). In reality it seems to be a bit of both - there is some general intelligence based on having been "trained on the internet", but it seems these super-human math/etc skills are very much from them having focused on training on those.
Francois Chollet mentioned that the test tries to avoid curve fitting (which he states is the main ability of LLMs). However, they specifically restricted the number of examples to do this. It is not beyond the realms of possibility that many examples could have been generated by hand though, and that the curve fitting has been achieved, rather than discrete programming.
Anyway, it’s all supposition. It’s difficult to know how genuine the results is, without knowledge of how it was actually achieved.