A friend is working for a newspaper. He records interviews. We tried all the software we could find to turn the recording (Dutch) into text but there is nothing that gives a helpful result. I know that a recording-to-text is different than speech-to-text but even when I use OK Google most of the time the results are horrible. So after all those years I am still a little skeptical.
Just one example from the preface to Chollet's "Deep Learning with Python":
> If you’ve picked up this book, you’re probably aware of the extraordinary progress that deep learning has represented for the field of artificial intelligence in the recent past. In a mere five years, we’ve gone from near-unusable image recognition and speech transcription, to superhuman performance on these tasks.
Come on, speech to text is still far from usable unless in a very limited scenarios. Why pretend it's different?