> The examples he gives are technologies where you can point the AI to a problem and trust the results will work as intended.
No, they're technologies where AI's have been tuned and tested and validated for there domain so that we after that process know that they will work as intended sufficiently often to measurably be a net benefit, just as you yourself go on to describe. They'll still make mistakes, just like you yourself point out. But they've been tested to ensure they make few enough mistakes to be worthwhile.
> What is the equivalent for LLMs?
The equivalent for LLMs is exactly the same process. To quote myself:
> he'd have made a far better point if he argued that LLMs also need to be evaluated carefully on a case by case basis like these other systems, rather than be taken on trust.
In other words: Test them on your use case, and validate their performance on that use case. Don't assume. We wouldn't take the performance of any domain specific model on trust, and there's no more reason to take the performance of an LLM on trust for a domain we've not tested it thoroughly for, and quite possibly fine tuned it for.
The difference is that LLMs do well enough to convince some people of the idea they can skip the testing and validation step, and that point - that people are prone to give them a level of trust that they should not be given - is valid. But extending that to dismissing them as "bullshit generators" as he did is equally ridiculous.