Like every model I've seen there is something like this:
>>A decoder is trained to predict the corresponding text...
Prediction of expected text in the context of the previous text.
While this is valuable in casual transcription, it can be extremely dangerous in serious contexts.
From personal experience, having given a deposition with an "AI" transcription, it will literally reverse the meanings of sentences.
This is because it produces the EXPECTED output in a context, and NOT THE ACTUAL OUTPUT.
Like a speaker that clips the output, these types of systems 'clip' the really valuable information out of a transcription. Worse yet, this is a completely silent failure, as the transcript LOOKS really good.
Basic info theory shows that there is more information contained in 'surprising' chunks of data than in expected ones. These systems actively work to substitute 'expected' speech to overwrite 'surprising' speech.
The transcript I got was utter trash, multiple pages of errata I had to submit when the normal is a couple of lines. And as I said, some literally reversed the meaning in a consequential way, and yet completely silently.
This kind of silent active failure mode is terrifying. Unless it is solved, and I see no way to solve it without removing ALL predictive algos from the system, these types of systems must not be used in any situation of serious consequence, at least not without real redundancy and backup.