Earlier quoted context omitted.
And the German is “subtitles of [public broadcaster] for [content network], 2017 I'm not sure this is really overfitting, the network does exactly what the training data demands. According to the training data silence art the end transcribes to a copyright notice or subtitle credits
fitting on noise in the training data is exactly what overfitting is. underfitting is smoothing out signal
Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
111–120 of 354 posts
Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#112Title should be changed to "OpenAI publishes evidence they trained on pirated movies".
How is this evidence of that fact? Honest question. I can see how this might show that subtitles from online sub communities are used, or that maybe even original subtitles from e.g. DVDs are used. But isn't it already known and admitted (and allowed?) that AI uses all sorts of copyrighted material to train models?
Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#113Earlier quoted context omitted.
That's the magic of money. Download your favorite artist's discography for personal use? If the MPAA had its way (and it occasionally has), torrenting that could bankrupt you. The AI industry - soaking up every bit of media available online for commercial purposes, often reproducing it nearly identically - has enough money and capital to influence things its way. And only its way, in case anyone was hoping this might…
> Download your favorite artist's discography for personal use? If the MPAA had its way (and it occasionally has), torrenting that could bankrupt you. I don't think that there are any clear examples of cases where ONLY downloading has resulted in huge fines. All the big bankrupting level fines have been for both downloading and sharing. You mention that 'torrenting' could bankrupt you, and that is true, but the main…
They [1, and others] been hunting and fining downloaders for over a decade now, with the only "evidence" being IP addresses connected with the torrent [2].
1: https://www.njordlaw.com/filesharing-and-downloading-films/q...
2: https://admin.ovpn.com/en/blog/online-integrity-new-threats-...
Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#114In English, silence is transcribed to "Please like and subscribe"
Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#115Earlier quoted context omitted.
ُThe Arabic text is the translator's self credit "Translated by Nancy Qanfar"
And the German is “subtitles of [public broadcaster] for [content network], 2017 I'm not sure this is really overfitting, the network does exactly what the training data demands. According to the training data silence art the end transcribes to a copyright notice or subtitle credits
Side-note: it's also yet more evidence that AI companies hoover all data with no regard for legality or copyright status, the very same offences that got other people in jail or with heavy fines.
Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#116Interesting! I used whipser last year to attempt to build an audio transcription tool but gave up due to excessive amount of hallucinated output no matter what model I used. It would produce seemingly ok output until you started paying attention. One example, it insisted that Biggie Smalls sings "Puttin five carrots in my baby girl ear". (its "carats"). It's apparently not useful in transcription as it don't reason […
Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#117Title should be changed to "OpenAI publishes evidence they trained on pirated movies".
How is this evidence of that fact? Honest question. I can see how this might show that subtitles from online sub communities are used, or that maybe even original subtitles from e.g. DVDs are used. But isn't it already known and admitted (and allowed?) that AI uses all sorts of copyrighted material to train models?
Indeed, the captioning is copyrighted work and you are not legally allowed to copy and redistribute it.
> But isn't it already known and admitted (and allowed?)
No, and I don't see where you got that from. Meta [1], OpenAI [2] and everybody else is being sued as we speak.
1: https://petapixel.com/2025/01/10/lawsuit-alleges-mark-zucker...
2: https://www.reuters.com/legal/litigation/openai-hit-with-new...
Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#118The same happens with whisper-large-v3 on Chinese transcription: silence is transcribed to something like "please upvote, share and favourite this video". I suspect they trained the model on some random YouTube video without carefully picking really useful data.
Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#119Earlier quoted context omitted.
How is this evidence of that fact? Honest question. I can see how this might show that subtitles from online sub communities are used, or that maybe even original subtitles from e.g. DVDs are used. But isn't it already known and admitted (and allowed?) that AI uses all sorts of copyrighted material to train models?
> I can see how this might show that subtitles from online sub communities are used, or that maybe even original subtitles from e.g. DVDs are used. Indeed, the captioning is copyrighted work and you are not legally allowed to copy and redistribute it. > But isn't it already known and admitted (and allowed?) No, and I don't see where you got that from. Meta [1], OpenAI [2] and everybody else is being sued as we speak.…
Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#120I suspected as others mentioned, these were extracted from torrents movies.