Earlier quoted context omitted.
In Czech, Whisper usually transcribes music as "Titulky vytvořil JohnyX" ("subtitles made by JohnyX") for the same reason.
Haha, trained on torrented movies! :-D The MPA must be so proud.
Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
31–40 of 354 posts
Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#32violets are blue
unregistered hypercam 2
Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#33Classic overfitting It's the LLM equivalent of thinking that an out-of-office reply is the translation: https://www.theguardian.com/theguardian/2008/nov/01/5
How is this overfitting, rather than a data quality / classification issue?
"Translated by Nancy Qanfar"
Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#34Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#35Earlier quoted context omitted.
> I wonder if the ZDF gave its approval for it being used for LLM training though? I am pretty sure they didn't get asked.
Just like the people forced to pay for ZDF under threat of imprisonment.
[1] https://en.wikipedia.org/wiki/ARD_ZDF_Deutschlandradio_Beitr...
Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#36"[ sub by sk cn2 ]"
or
"Anyways, thanks for watching! Please subscribe and like! Thanks for watching! Bye!"
or
"This is the end of the video. Thank you for watching. If you enjoyed this video, please subscribe to the channel. Thank you."
Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#37The same happens with whisper-large-v3 on Chinese transcription: silence is transcribed to something like "please upvote, share and favourite this video". I suspect they trained the model on some random YouTube video without carefully picking really useful data.
The videos I tried to transcribe were also Mandarin Chinese, using whisper-large-v3. Besides the usual complaints that it would phonetically "mishear" things and generate nonsense, it was still surprisingly good, compared to other software I played around with.
That said, it would often invent names for the speakers and prefix their lines, or randomly switch between simplified and traditional Chinese. For the videos I tested, intermittent silence would often result in repeating the last line several times, or occasionally, it would insert direction cues (in English for some reason). I've never seen credits or anything like that.
In one video I transcribed, somebody had a cold and was sniffling. Whisper decided the person was crying (transcribed as "* crying *", a cough was turned into "* door closing *"). It then transcribed the next line as something quite unfriendly. It didn't do that anymore after I cut the sniffling out (but then the output switched back to traditional Chinese again).
Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#38I wonder if hallucinated copyright claims (esp. like the ZDF one at the bottom of the OP) will be introduced as evidence in one of the court cases against "big AI"
Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#39leaving personal comments, jokes, reactions, intros in subtitles is very common in eastern cultures.
Turkish readers will probably remember “esekadam iyi seyirler diler” :)
Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#40Earlier quoted context omitted.
> I wonder if the ZDF gave its approval for it being used for LLM training though? I am pretty sure they didn't get asked.
Just like the people forced to pay for ZDF under threat of imprisonment.