Aside from the S-versus-exp issue, this area is one of these things where there's a kind of disconnect between my personal professional experience with LLMs and the criteria measures he's talking about. LLMs to me have this kind of superficially impressive feel where it seems impressive in its capabilities, but where, when it fails, it fails dramatically, in a way humans never would, and it never gets anywhere near what's necessary to actually be helpful on finishing tasks, beyond being some kind of gestalt template or prototype.
I feel as if there needs to be a lot more scrutiny on the types of evaluation tasks being provided — whether they are actually representative of real-world demands, or if they are making them easy to look good, and also more focus on the types of failures. Looking through some of the evaluation tasks he links to I'm more familiar with, they seem kind of basic? So not achieving parity with human performance is more significant than it seems. I also wonder, in some kind of maxmin sense, whether we need to start focusing more on worst-case failure performance rather than best-case goal performance.
LLMs are really amazing in some sense, and maybe this essay makes some points that are important to keep in mind as possibilities, but my general impression after reading it is it's kind of missing the core substance of AI bubble claims at the moment.