Classic overfitting It's the LLM equivalent of thinking that an out-of-office reply is the translation: https://www.theguardian.com/theguardian/2008/nov/01/5
How is this overfitting, rather than a data quality / classification issue?
Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
41–50 of 354 posts
Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#42Earlier quoted context omitted.
Looks like they used more official sources for German - there, silence is apparently hallucinated as "Untertitelung des ZDF für funk, 2017" according to one of the comments on the issue. Which makes sense, as the public broadcasters' "Mediathek" is probably the largest freely available resource of subtitled videos in Germany. I wonder if the ZDF gave its approval for it being used for LLM training though?
Most content from Funk (youtubers funded by public german broadcasters) is available on youtube without any geoblocking or other limitations.
Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#43The same happens with whisper-large-v3 on Chinese transcription: silence is transcribed to something like "please upvote, share and favourite this video". I suspect they trained the model on some random YouTube video without carefully picking really useful data.
Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#44Earlier quoted context omitted.
Looks like they used more official sources for German - there, silence is apparently hallucinated as "Untertitelung des ZDF für funk, 2017" according to one of the comments on the issue. Which makes sense, as the public broadcasters' "Mediathek" is probably the largest freely available resource of subtitled videos in Germany. I wonder if the ZDF gave its approval for it being used for LLM training though?
Most content from Funk (youtubers funded by public german broadcasters) is available on youtube without any geoblocking or other limitations.
Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#45Earlier quoted context omitted.
Haha, trained on torrented movies! :-D The MPA must be so proud.
It's absolutely insane that these companies can't be held liable for what is obvious piracy.
The AI industry - soaking up every bit of media available online for commercial purposes, often reproducing it nearly identically - has enough money and capital to influence things its way. And only its way, in case anyone was hoping this might change anything at all for the little guy.
Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#46roses are red violets are blue unregistered hypercam 2
Silence is golden,
Translated by Nancy,
To copyright, we aren't beholden
Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#47Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#48Earlier quoted context omitted.
How is this overfitting, rather than a data quality / classification issue?
ُThe Arabic text is the translator's self credit "Translated by Nancy Qanfar"
I'm not sure this is really overfitting, the network does exactly what the training data demands. According to the training data silence art the end transcribes to a copyright notice or subtitle credits
Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#49Earlier quoted context omitted.
It's absolutely insane that these companies can't be held liable for what is obvious piracy.
That's the magic of money. Download your favorite artist's discography for personal use? If the MPAA had its way (and it occasionally has), torrenting that could bankrupt you. The AI industry - soaking up every bit of media available online for commercial purposes, often reproducing it nearly identically - has enough money and capital to influence things its way. And only its way, in case anyone was hoping this might…
Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#50I wonder if hallucinated copyright claims (esp. like the ZDF one at the bottom of the OP) will be introduced as evidence in one of the court cases against "big AI"
It already has been and meta won the lawsuit because corporations are sacrosanct.