I am really confused on how it can saturate ARC-AGI but still perform poorly on aggregated benchmarks: https://artificialanalysis.ai/models Perhaps if it was allowed this custom harness for all benchmarks it would similarily saturate?
And, you know, maybe also some funny business. I think it's good to be a little suspicious of a model that happens to shoot upwards in performance on a specific benchmark while also kind of keeping up with the pack on a bunch of other benchmarks.