There's an art to subtitling that goes beyond mere speech-to-text processing. Sometimes it's better to paraphrase dialog to reduce the amount of text that needs to be read. Sometimes you need to name a voice as unknown, to avoid spoilers. Sometimes the positioning on the screen matters. I hope the model can be made to understand all this.
> better to paraphrase dialog to reduce the amount of text that needs to be read. That's just bad destructive art, especially for a foreign language that you partially know. > Sometimes you need to name a voice as unknown, to avoid spoilers. Don't name any, that's what your own eye-ear voice recognition/matching and positioning are for (also reduces the amount of text) > Sometimes the positioning on the screen matter…
That’s tricky when one or more speakers aren’t visible.