The most interesting part, even more than ARC 3 score, to me is that this is the first model I recall seeing that scores lower on Max than High reasoning effort on some coding benchmarks: Terminal-Bench 4.0: High (57.9%), Max (56.7%) DeepSWE: High (73.3%), Max (71.5%) It _loses_ 1-2% performance going to High from Max
That's quite common with many models, after "High" reasoning, over-thinking starts occurring and the model skips over the right solution by convincing itself otherwise.
Such as?
I can't think of any. Diminishing returns, yes. Occasionally flat, yes. Downright regression, no.