So Google hasn't used an LLM to generate and test weird queries ? This is not putting the bar very high for the whole industry... There'd be so much to gain from a clean deployment... Either it hard, either it is a rush. As a machine learnist, I believe it's actually impossible, by design of the autoregressive LLM. This race may we'll be partially to the bottom.
You don't even need an LLM for that. Google will almost certainly have tested.
The test result is just politically-unacceptable within the company: It doesn't work, it's a architectural issue inherent to the technology, we can't fix it.
Instead, they just rush to patch any specific, individual errors that show up, and claim that these errors are "rare exceptions" or "never happened".
What's going on here is that Google (and most other AI firms) are just trying to gaslight the world about how error-prone AI is, because they're in too deep and can't accept the reality themselves.