Earlier quoted context omitted.
Are we reading the same chart? They have Sonnet You have to test each task obviously but it is not a bad model on its face.
They have updated it
> Anthropic did post an official explanation, stating the original chart used a "simpler methodology" that "underestimated Sonnet 5's performance." The new chart supposedly uses their "standard methodology."
Oops!