The claim isn't so wild when it's a generalist versus a finetune trained specifically on the tasks being benchmarked.