Earlier quoted context omitted.
Indeed, with another model I would get persistent transcriptions of silent parts into 'Thanks for watching!' or '[MUSIC]'. Pretty dumb that this failure mode wasn't caught in some QA process, and there are now multiple transcription models suffering from the same issue. Having silent parts in your input audio seems like it should be a very common occurrence...
whisper MUST be combined with silence detection / VAD
What good is a speech recognition tool that literally hears imaginary voices?