Live data from Hacker News

Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

github.com

41–50 of 354 posts

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#41

Classic overfitting It's the LLM equivalent of thinking that an out-of-office reply is the translation: https://www.theguardian.com/theguardian/2008/nov/01/5

How is this overfitting, rather than a data quality / classification issue?

It is a data quality issue which caused the model to overfit.

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#42
post #11

Earlier quoted context omitted.

Looks like they used more official sources for German - there, silence is apparently hallucinated as "Untertitelung des ZDF für funk, 2017" according to one of the comments on the issue. Which makes sense, as the public broadcasters' "Mediathek" is probably the largest freely available resource of subtitled videos in Germany. I wonder if the ZDF gave its approval for it being used for LLM training though?

Most content from Funk (youtubers funded by public german broadcasters) is available on youtube without any geoblocking or other limitations.

Ah, ok, thanks for the info, TIL! "We are funk – the first public service content network that started on October 1, 2016. We create online-only content on social networks and third-party platforms, including YouTube, Instagram, Snapchat, TikTok, Spotify, Apple Music or Twitch for 14-29 year-olds." (https://presse.funk.net/das-ist-funk/, scroll down for the English version). I live in Germany, and I even watch public broadcasters regularly, but this is the first time I have heard about funk (I even initially thought it was misspelled, usually it's written with a capital F). But I'm not part of the targeted audience (not now, nor even back in 2016 when it was launched), so all good...

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#43
post #24

The same happens with whisper-large-v3 on Chinese transcription: silence is transcribed to something like "please upvote, share and favourite this video". I suspect they trained the model on some random YouTube video without carefully picking really useful data.

Indeed, with another model I would get persistent transcriptions of silent parts into 'Thanks for watching!' or '[MUSIC]'. Pretty dumb that this failure mode wasn't caught in some QA process, and there are now multiple transcription models suffering from the same issue. Having silent parts in your input audio seems like it should be a very common occurrence...

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#44
post #11

Earlier quoted context omitted.

Looks like they used more official sources for German - there, silence is apparently hallucinated as "Untertitelung des ZDF für funk, 2017" according to one of the comments on the issue. Which makes sense, as the public broadcasters' "Mediathek" is probably the largest freely available resource of subtitled videos in Germany. I wonder if the ZDF gave its approval for it being used for LLM training though?

Most content from Funk (youtubers funded by public german broadcasters) is available on youtube without any geoblocking or other limitations.

I’m pretty sure that content doesn’t come with a license granting unlimited usage rights.

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#45

Earlier quoted context omitted.

Haha, trained on torrented movies! :-D The MPA must be so proud.

It's absolutely insane that these companies can't be held liable for what is obvious piracy.

That's the magic of money. Download your favorite artist's discography for personal use? If the MPAA had its way (and it occasionally has), torrenting that could bankrupt you.

The AI industry - soaking up every bit of media available online for commercial purposes, often reproducing it nearly identically - has enough money and capital to influence things its way. And only its way, in case anyone was hoping this might change anything at all for the little guy.

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#47
In Russian it often hallucinates "Субтитры сделал DimaTorzok" ("Subtitles by DimaTorzok") at the end of things. Interestingly, I wasn't able to find any YouTube videos with that name in the subtitles, so it's not like it's in a lot of training data.

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#48
post #33

Earlier quoted context omitted.

How is this overfitting, rather than a data quality / classification issue?

ُThe Arabic text is the translator's self credit "Translated by Nancy Qanfar"

And the German is “subtitles of [public broadcaster] for [content network], 2017

I'm not sure this is really overfitting, the network does exactly what the training data demands. According to the training data silence art the end transcribes to a copyright notice or subtitle credits

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#49
post #45

Earlier quoted context omitted.

It's absolutely insane that these companies can't be held liable for what is obvious piracy.

That's the magic of money. Download your favorite artist's discography for personal use? If the MPAA had its way (and it occasionally has), torrenting that could bankrupt you. The AI industry - soaking up every bit of media available online for commercial purposes, often reproducing it nearly identically - has enough money and capital to influence things its way. And only its way, in case anyone was hoping this might…

The movie industry also has some money and lobbying power. Surely this is a way larger threat than any single torrenter could ever be?

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#50

I wonder if hallucinated copyright claims (esp. like the ZDF one at the bottom of the OP) will be introduced as evidence in one of the court cases against "big AI"

It already has been and meta won the lawsuit because corporations are sacrosanct.

Do you mean that specifically a hulucibated text "copyright by not-meta" made it into evidence? Or are you talking about copyright generally?
Post reply on HN