Earlier quoted context omitted.
This also works; I upvoted your comment.
I have discovered a truly marvelous proof of how to smash that like and subscribe button, which this comment box is too small to contain.
Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
101–110 of 354 posts
Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#102Earlier quoted context omitted.
The movie industry also has some money and lobbying power. Surely this is a way larger threat than any single torrenter could ever be?
The fact that this is propping up the entire AI industry adds additional weight. When legislating or deciding court cases, some won't be willing to pop the cash cow, some will be worried about falling behind countries that don't enforce copyright evenly. IP owners are trying to go after the AI industry, with only mixed to poor success.
Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#103Earlier quoted context omitted.
That's the magic of money. Download your favorite artist's discography for personal use? If the MPAA had its way (and it occasionally has), torrenting that could bankrupt you. The AI industry - soaking up every bit of media available online for commercial purposes, often reproducing it nearly identically - has enough money and capital to influence things its way. And only its way, in case anyone was hoping this might…
> Download your favorite artist's discography for personal use? If the MPAA had its way (and it occasionally has), torrenting that could bankrupt you. I don't think that there are any clear examples of cases where ONLY downloading has resulted in huge fines. All the big bankrupting level fines have been for both downloading and sharing. You mention that 'torrenting' could bankrupt you, and that is true, but the main…
Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#104Earlier quoted context omitted.
When I was taught mathematics, the zero value was always considered the most important edge case. You prove something for N=0 (or N=1), then for N=M+1. It's even more important in audio DSP: processing near-zeroes can end up being extremely CPU intensive, look up denormal/subnormal floats.
Yeah, I studied mathematics (algebra and number theory) and zero is the point, often sporting discontinuities, or weird asymptotic behavior. Quite a lot of algorithms use some form of division and zero is the only number in our typical structures (Z, Q, R, C), that cannot be used to divide with.
Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#105Who is Nicolai Winther? https://medium.com/@lehandreassen/who-is-nicolai-winther-985...
Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#106It would produce seemingly ok output until you started paying attention.
One example, it insisted that Biggie Smalls sings "Puttin five carrots in my baby girl ear". (its "carats").
It's apparently not useful in transcription as it don't reason [sic].
Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#107Title should be changed to "OpenAI publishes evidence they trained on pirated movies".
Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#108I wonder if hallucinated copyright claims (esp. like the ZDF one at the bottom of the OP) will be introduced as evidence in one of the court cases against "big AI"
Evidence against what? "Big AI" is transparent and open about the fact they use all sorts of copyrighted material to train the data. How would "we see an exact chunk of text from our copyrighted material" add to that?
So not only are they training on copyrighted material, but they didn't even pay for it once, and then they didn't even do minimal data cleaning before training. Which, by the way, is the type of cleaning their LLMs could have done.
Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#109Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#110The same happens with whisper-large-v3 on Chinese transcription: silence is transcribed to something like "please upvote, share and favourite this video". I suspect they trained the model on some random YouTube video without carefully picking really useful data.