Earlier quoted context omitted.
I really doubt LLM benchmarks are reflective of real world user experience ever since they claimed GPT-4o hallucinated less than the original GPT-4.
I begin to believe LLM benchmarks are like european car mileage specs. They say its 4 Liter / 100km but everyone knows it's at least 30% off (same with WLTP for EVs).
You need to remove your shoe and drive with like two toes to get the speed just right, though.
Test drivers I have done this with takes off their shoes or use ballerina shoes.