The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…
What does "AGI" or "effective AGI" even mean, and why should anyone even care whether this unclear thing has been "reached" or not? Computer chips got faster, but 2026 edition. Why the artificial ceiling/category/goal labelled "AGI"? I'd much rather like to talk about what this enables, instead of discussing whether a category someone made up applies here or not.
> AGI is essentially the equivalent of a median human that could be hired as a remote co-worker... capable of performing any task that one would be satisfied with a remote colleague doing via a computer.
So... unless you hear of a company replacing their workforce with OpenAI agents, I don't think we're there yet.