Live data from Hacker News

Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

github.com

21–30 of 354 posts

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#21
post #3

to save you a lookup: The Arabic text "رجمة نانسي قنقر" translates to English as: "Nancy Qanqar's translation" or "Translation by Nancy Qanqar" "رجمة" means "translation" and "نانسي قنقر" is the name "Nancy Qanqar"

And it seems to be because the training data is largely unofficial subtitles from movies. Which often have a string like "Translated by X" at the end of the movie which is often silent while credits roll.

I'm sure they totally did not pirate the audio of said movies.

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#22
post #11

Earlier quoted context omitted.

Looks like they used more official sources for German - there, silence is apparently hallucinated as "Untertitelung des ZDF für funk, 2017" according to one of the comments on the issue. Which makes sense, as the public broadcasters' "Mediathek" is probably the largest freely available resource of subtitled videos in Germany. I wonder if the ZDF gave its approval for it being used for LLM training though?

> I wonder if the ZDF gave its approval for it being used for LLM training though? I am pretty sure they didn't get asked.

Just like the people forced to pay for ZDF under threat of imprisonment.

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#23
post #3

to save you a lookup: The Arabic text "رجمة نانسي قنقر" translates to English as: "Nancy Qanqar's translation" or "Translation by Nancy Qanqar" "رجمة" means "translation" and "نانسي قنقر" is the name "Nancy Qanqar"

You've got a little typo, it's not "رجمة", it's "ترجمة" that means translation, the ت at the beginning is missing.

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#24
The same happens with whisper-large-v3 on Chinese transcription: silence is transcribed to something like "please upvote, share and favourite this video". I suspect they trained the model on some random YouTube video without carefully picking really useful data.

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#25
post #11

Earlier quoted context omitted.

And it seems to be because the training data is largely unofficial subtitles from movies. Which often have a string like "Translated by X" at the end of the movie which is often silent while credits roll.

Looks like they used more official sources for German - there, silence is apparently hallucinated as "Untertitelung des ZDF für funk, 2017" according to one of the comments on the issue. Which makes sense, as the public broadcasters' "Mediathek" is probably the largest freely available resource of subtitled videos in Germany. I wonder if the ZDF gave its approval for it being used for LLM training though?

Most content from Funk (youtubers funded by public german broadcasters) is available on youtube without any geoblocking or other limitations.

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#27

In Italian as well there are random hallucination when parsing silence, something like: “Thank you for watching”, “Subtitles by…”

I wouldn't be surprised if "like share and subscribe" also shows up at some point.

"наша зброя в цей момент -- вподобайка і комент"

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#28
post #8
post #3

to save you a lookup: The Arabic text "رجمة نانسي قنقر" translates to English as: "Nancy Qanqar's translation" or "Translation by Nancy Qanqar" "رجمة" means "translation" and "نانسي قنقر" is the name "Nancy Qanqar"

In Czech, Whisper usually transcribes music as "Titulky vytvořil JohnyX" ("subtitles made by JohnyX") for the same reason.

Haha, trained on torrented movies! :-D

The MPA must be so proud.

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#30

Classic overfitting It's the LLM equivalent of thinking that an out-of-office reply is the translation: https://www.theguardian.com/theguardian/2008/nov/01/5

How is this overfitting, rather than a data quality / classification issue?
Post reply on HN