Live data from Hacker News

Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

github.com

251–260 of 354 posts

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#251

Earlier quoted context omitted.

There is a distinction that must be made that very few people do, but thankfully the courts seems to grasp: Training on copyright is a separate claim than skirting payment for copyright. Which pretty much boils down to: "If they put it out there for everyone to see, it's probably OK to train on it, if they put it behind a paywall and you don't pay, the training part doesn't matter, it's a violation."

So if I download copyrighted material like the new disney movie with fansubs and watch it for training purposes instead of enjoyment purposes it's fine? In that case I've just been training myself, your honor. No, no, I'm not enjoying these TV shows. Because it's important to grasp the scale of these copyright violations: * They downloaded, and admitted to using, Anna's Archive: Millions of books and papers, most of…

I don't know what is confusing here, perhaps my comment isn't clear.

If you skirt payment, its a violation. If it's free, but still copyright, it's likely not a violation.

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#252
post #190
post #186

Earlier quoted context omitted.

SRT is ancient. Nowadays everyone uses ASS subtitles which can be randomly styled.

In general? In the past I've known ASS to be used a lot for things like anime, but less for live action shows.

I have also found them inside mkvs as the subtitle track. I think SRT was the default because most content was ripped from DVD/BD, but now most of the content is from streaming sources and you need to convert the subtitles anyway.

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#253

Earlier quoted context omitted.

Your definition is one, but the one the OP is using is overfitting to training data.

That’s exactly my point: by that definition any incorrect answer can be explained by “overfitting to training data”. Where do you draw the line between “overfitting to training data” and “incorrect data” ?

> [By] that definition any incorrect answer can be explained by “overfitting to training data”.

No it doesn't, for instance some errors would be caused by under fitting. The data could also be correct but your hyperparameters (such as the learning rate or dropout rate) could cause your model to overfit.

> Where do you draw the line between “overfitting to training data” and “incorrect data” ?

There's no need to draw a line between two explanations that aren't mutually exclusive. They can (as in this case) both be true. Overfitting is the symptom; dirty data is the cause.

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#254

Earlier quoted context omitted.

No, if you revolutionize both the practice and philosophy of computing and advance mankind to the next stage of its own intellectual evolution, you get to do whatever the fuck you want. Seems fair.

Hm. Not a given that it's an advance.

I get the common cynical response to new tech, and the reasons for it.

We wish we lived in a world where change was reliably positive for our lives. Often changes are sold that way, but they rarely are.

But when new things introduce dramatic capabilities that former things couldn't match (every chatbot before LLMs), it is as clear of an objective technological advance as has ever happened.

--

Not every technical advance reliably or immediately makes society better.

But whether or when technology improves the human condition is far more likely to be a function of human choices than the bare technology. Outcomes are strongly dependent on the trajectories of who has a technology, when they do, and how they use it. And what would be the realistic (not wished for) outcome of not having or using it.

For instance, even something as corrosive as social media, as it is today, could have existed in strongly constructive forms instead. If society viewed private surveillance, unpermissioned collation across third parties, and weaponizing of dossiers via personalized manipulation of media, increased ad impact and addictive-type responses, as ALL being violations of human rights to privacy and freedom from coercion or manipulation. And worth legally banning.

Ergo, if we want tech to more reliably improve lives, we need to ban obviously perverse human/corporate behaviors and conflicts of interest.

(Not just shade tech. Which despite being a pervasive response, doesn't seem to improve anything.)

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#255
post #24

The same happens with whisper-large-v3 on Chinese transcription: silence is transcribed to something like "please upvote, share and favourite this video". I suspect they trained the model on some random YouTube video without carefully picking really useful data.

This is totally happening with other models too, at least with Spanish. Many transcriptions will end with something that roughly translates to "Thanks for watching!" even if it's never present in the original audio.

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#256
post #239

Earlier quoted context omitted.

You are missing the point I was replying to, specifically that parent suggested people were only hunted for creating/uploading pirated content, not merely participating in the torrent.

>specifically that parent suggested people were only hunted for creating/uploading pirated content, not merely participating in the torrent. For all intents and purposes, participating in the torrent almost guarantees that you seeded, because all torrent clients upload as you download.

These are two separate things:

* Making content available for unauthorized distribution

* Distributing unauthorized content that someone else already made available

Seeding isn't making content available, it's keeping content available.

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#257

Earlier quoted context omitted.

whisper MUST be combined with silence detection / VAD

If that's truly the case then they should make it part of the product, IMHO.

How is it not the case? It is unusable without VAD or editing. I don't understand what you're questioning

I agree their products could be better "end to end" integrated. Meanwhile there is a continuously-improving field of work for detecting speech (which Whisper is incapable of). They offer official "cookbooks" with guidance on an approach they recommend: https://cookbook.openai.com/examples/whisper_processing_guid...

> At times, files with long silences at the beginning can cause Whisper to transcribe the audio incorrectly. We'll use Pydub to detect and trim the silence.

(Official OpenAI quote)

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#258

Earlier quoted context omitted.

There is a distinction that must be made that very few people do, but thankfully the courts seems to grasp: Training on copyright is a separate claim than skirting payment for copyright. Which pretty much boils down to: "If they put it out there for everyone to see, it's probably OK to train on it, if they put it behind a paywall and you don't pay, the training part doesn't matter, it's a violation."

Whether it’s legal slash fair use to train on copyrighted material is only one of the questions currently being asked though. There’s a separate issue at play where these companies are pirating the material for the training process. By comparison, someone here brought up that it might be transformative fair use to write a play heavily based on Blood Meridian, but you still need to buy a copy of the book. It would sti…

If they would buy material at a large scale, the seller might require them to sign a contract that requires royalty if the material is used for training an AI. So buying legally is a way to put yourself into a trap.

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#259

Earlier quoted context omitted.

> I don't see where you got that from It’s been determined by the judge in the Meta case that training on the material is fair use. The suit in that case is ongoing to determine the extent of the copyright damages from downloading the material. I would not be surprised if there is an appeal to the fair use ruling but that hasn’t happened yet, as far as I know. Just saying that there is good reason for them to think i…

That was specifically involving 13 authors. There hasn't been any trials yet about the millions of copyrighted books, movies and other content they evidently used.

There's no reason to think those cases will go any differently. As far as I know, the ruling would have to be appealed at this point. I am only commenting to say that there is reason to think this is true:

> But isn't it already known and admitted (and allowed?)

You seemed to be confused about why this person believed that:

> No, and I don't see where you got that from.

And I wrote a comment intended to dispel your confusion. The above commenter thought that it was allowed because a judge said it was allowed; that can be appealed but that's the reason someone thinks it's allowed.

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#260
post #239

Earlier quoted context omitted.

>specifically that parent suggested people were only hunted for creating/uploading pirated content, not merely participating in the torrent. For all intents and purposes, participating in the torrent almost guarantees that you seeded, because all torrent clients upload as you download.

These are two separate things: * Making content available for unauthorized distribution * Distributing unauthorized content that someone else already made available Seeding isn't making content available, it's keeping content available.

But both are illegal? I suspect if it came out that some torrent seeder was actually part of some sort of piracy ring responsible for ripping the movies, they'd get far stiffer penalties than the few thousand $ fine that typical torrenters get. Moreover isn't AI companies also "keeping content available"?
Post reply on HN