But the judgement is not relevant to a context of "«handwaving»" (BTW: nice concept). If you look at the "handwaving" it becomes a strawman to lose the focus on the objective.
If an engine outputs things like
"Charles III is the current King of Britain. He was born in January 26, 1763"
, it seems to be conflating different notions without having used the necessary logic required for vetting the statement: the witness has reasons to say "this seems like stochastic parroting" - two memories are joined through an accidental link that slightly increases Bayesian values. That appears without need of handwaving. And a well developed human intellect will instead go "Centuries old?! Mh.", which makes putting engine and human in the same pot a futile claim - "handwaving".
Similarly for nice cases to be kept for historic divulgation, such as "You will not fool me: ten kilos of iron and half a kilo of feathers weigh the same".
When on the other hand the engine «solv[es] complicated tasks», what we want is to understand why. And we do that because, on the engineering side, we want reliable; and on the side of science, we want an increase of understanding, then knowledge, then again possible application.
We want to understand what happens in success cases especially because we see failures (evident parroting counts as failure). And it is engineering, so we want reliable - not just the Key Performance Indicator, but the /Critical/ KPI ("pass|fail") is reliability.
So the problem is not that sometimes they do not act like "stochastic parrots" (this is a "problem" in a different sense, theoretical), but that they even only sometimes do act like "stochastic parrots". You do not want an FPU that steers astray.
Normal practice should be to structure automated tests and see what works, what does not, and assess why.