It significantly improves upon GPT-4o on my Extended NYT Connections Benchmark. 22.4 -> 33.7 ( https://github.com/lechmazur/nyt-connections ).
The answers are certainly in the training set, likely many times over.
I’d be curious to see performance on Bracket City, which was featured here on HN yesterday.