The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…
I genuinely don’t understand how an adult can say this with a straight face. I can take any single of my hobbies, start a mildly advanced conversation with Fable about the hobby and, within 5-10 turns, get it to contradict itself about something fundamental, lie or give bad or dangerous advice.