Live data from Hacker News

Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

github.com

181–190 of 354 posts

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#181
post #99

Earlier quoted context omitted.

In Chinese, it always added something like "For study/research purpose only. Please delete after 48 hours." This is what those volunteers added in subtitles of (pirated) movies/shows.

Fair, if AI companies are allowed to download pirated content for "learning", why ordinary people cannot.

There is so much damning evidence that AI companies have committed absolutely shocking amounts of piracy, yet nothing is being done.

It only highlights how the world really works. If you have money you get to do whatever the fuck you want. If you're just a normal person you get to spend years in jail or worse.

Reminds me of https://www.youtube.com/watch?v=8GptobqPsvg

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#182
post #86

Earlier quoted context omitted.

This person refers to the German television and radio fee (Rundfunkgebühren).[1] It is a state-mandated system that ensures free (as in free speech) and (relatively) neutral public broadcasting institutions. There is a constant and engaged discussion, because every household in Germany has to pay this fee. Exceptions are made only for low-income households. [1] https://en.wikipedia.org/wiki/ARD_ZDF_Deutschlandradio_B…

A constant discussion, lately fueled by extremist parties (AfD) who feel treated unfairly by (amongst others) the public broadcasters (which has parallels to Trump's recent campaign against public broadcasters in the US).

[deleted]

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#183
post #29

Earlier quoted context omitted.

I can clap with one hand (fingers on palm) and it produces a clapping sound.

Your brain merely hallucinates a clapping sound as "Translation by Nancy Qunqar" enters your ears.

Good one. :)

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#184
post #57

Earlier quoted context omitted.

How is this overfitting, rather than a data quality / classification issue?

Isn't overfitting just when the model picks up on an unintended pattern in the training data? Isn't that precisely what this is?

[deleted]

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#185
post #83
post #58

Title should be changed to "OpenAI publishes evidence they trained on pirated movies".

How is this evidence of that fact? Honest question. I can see how this might show that subtitles from online sub communities are used, or that maybe even original subtitles from e.g. DVDs are used. But isn't it already known and admitted (and allowed?) that AI uses all sorts of copyrighted material to train models?

> How is this evidence of that fact?

The contention is that the specific translated text appears largely from illegal translations (i.e., fansubs) and not from authorized translations. And from a legal perspective, that would basically mean there's no way they could legally have appropriated that material.

> But isn't it already known and admitted (and allowed?) that AI uses all sorts of copyrighted material to train models?

Technically, everything is copyrighted. But your question is really about permission. Some of the known corpuses for AI training include known pirate materials (e.g., libgen), but it's not known whether or not the AI companies are filtering out those materials from training. There's a large clutch of cases ongoing right now about whether or not AI training is fair use or not, and the ones that have resolved at this point have done so on technical grounds rather than answering the question at stake.

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#186
post #155

Earlier quoted context omitted.

That is not the case here - I never encountered this with whisper-large-v3 or similar ASR models. Part of the reason, I guess, is that those subs are burnt into the movie, which makes them hard to extract. Standalone subs need the corresponding video resource to match the audio and text. So nothing is better than YouTube videos which are already aligned.

At least for English, those "fansubs" aren't typically burnt into the movie*, but ride along in the video container (MP4/MKV) as subtitle streams. They can typically be extracted as SRT files (plain text with sentence level timestamps). *Although it used to be more common for AVI files in the olden days.

SRT is ancient. Nowadays everyone uses ASS subtitles which can be randomly styled.

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#187

Earlier quoted context omitted.

Fair, if AI companies are allowed to download pirated content for "learning", why ordinary people cannot.

There is so much damning evidence that AI companies have committed absolutely shocking amounts of piracy, yet nothing is being done. It only highlights how the world really works. If you have money you get to do whatever the fuck you want. If you're just a normal person you get to spend years in jail or worse. Reminds me of https://www.youtube.com/watch?v=8GptobqPsvg

If you owe the bank $1,000 you have a problem.

If you owe the bank $100,000,000 the bank has a problem.

We live in an era where the president of the United States uses his position to pump crypto scams purely for personal profit.

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#188

Earlier quoted context omitted.

> I can see how this might show that subtitles from online sub communities are used, or that maybe even original subtitles from e.g. DVDs are used. Indeed, the captioning is copyrighted work and you are not legally allowed to copy and redistribute it. > But isn't it already known and admitted (and allowed?) No, and I don't see where you got that from. Meta [1], OpenAI [2] and everybody else is being sued as we speak.…

> I don't see where you got that from It’s been determined by the judge in the Meta case that training on the material is fair use. The suit in that case is ongoing to determine the extent of the copyright damages from downloading the material. I would not be surprised if there is an appeal to the fair use ruling but that hasn’t happened yet, as far as I know. Just saying that there is good reason for them to think i…

That was specifically involving 13 authors.

There hasn't been any trials yet about the millions of copyrighted books, movies and other content they evidently used.

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#189
post #19

Little did you all know, this is just being mechanical turked by Nancy Qunqar. Way to go Nancy! Keep up the good work, ya crazy bastard!

Is this spam? That name only shows as an instagram account and this thread. If you pay for insta followers is this how they get them now? Haha

That’s the name in the Arabic text hallucinated by the model :)

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#190
post #186
post #155

Earlier quoted context omitted.

At least for English, those "fansubs" aren't typically burnt into the movie*, but ride along in the video container (MP4/MKV) as subtitle streams. They can typically be extracted as SRT files (plain text with sentence level timestamps). *Although it used to be more common for AVI files in the olden days.

SRT is ancient. Nowadays everyone uses ASS subtitles which can be randomly styled.

In general? In the past I've known ASS to be used a lot for things like anime, but less for live action shows.
Post reply on HN