Well now I know how I’m going to start filling awkward silences in meetings.
Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
151–160 of 354 posts
Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#152Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#153Same happened to me with English: I've got "Thanks for watching" many times.
Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#154Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#155Earlier quoted context omitted.
In Chinese, it always added something like "For study/research purpose only. Please delete after 48 hours." This is what those volunteers added in subtitles of (pirated) movies/shows.
That is not the case here - I never encountered this with whisper-large-v3 or similar ASR models. Part of the reason, I guess, is that those subs are burnt into the movie, which makes them hard to extract. Standalone subs need the corresponding video resource to match the audio and text. So nothing is better than YouTube videos which are already aligned.
*Although it used to be more common for AVI files in the olden days.
Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#156Earlier quoted context omitted.
It's absolutely insane that these companies can't be held liable for what is obvious piracy.
What's insane is copyright. How come you can own intellectual property but not pay a property tax? The ecosystem would be much healthier if to get copyright protections you should declare value of your IP (that you are obligated to sell for if the buyer pops up) and pay tax on this for every year you hold the IP.
Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#157Earlier quoted context omitted.
Your definition is one, but the one the OP is using is overfitting to training data.
That’s exactly my point: by that definition any incorrect answer can be explained by “overfitting to training data”. Where do you draw the line between “overfitting to training data” and “incorrect data” ?
Not really, getting 94381294*123=... wrong, but close within the actual answer, cannot be overfitting since it wasn't in the training data.
Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#158Earlier quoted context omitted.
Indeed, with another model I would get persistent transcriptions of silent parts into 'Thanks for watching!' or '[MUSIC]'. Pretty dumb that this failure mode wasn't caught in some QA process, and there are now multiple transcription models suffering from the same issue. Having silent parts in your input audio seems like it should be a very common occurrence...
whisper MUST be combined with silence detection / VAD
Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#159The same happens with whisper-large-v3 on Chinese transcription: silence is transcribed to something like "please upvote, share and favourite this video". I suspect they trained the model on some random YouTube video without carefully picking really useful data.
Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#160Earlier quoted context omitted.
How is this overfitting, rather than a data quality / classification issue?
ُThe Arabic text is the translator's self credit "Translated by Nancy Qanfar"