Capable enough LLMs are human level for lots of things. Reinforcement learning from ai feedback is a thing (the anthropic claude models use that). Strictly speaking, it's not necessary to have humans in the loop for a lot of these things. Some are hesitant to admit we've created human level general intelligence but saying otherwise doesn't really hold up to scrutiny.
I see people saying things like this but I have yet to see anyone show data for a non-trivial workflow with human-level accuracy over a wide range of inputs, without a human in the loop.
As a decent first-order metric - follow the ratio of companies getting money for using LLMs to do something, to companies getting money for providing LLMs and associated tooling to others. The bigger that ratio gets, the more real world impact LLMs are having.