Live data from Hacker News

Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

github.com

281–290 of 354 posts

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#281

Earlier quoted context omitted.

There is indeed, but not when you’re torrenting (i.e. you can’t download without also uploading).

Even when you are torrenting, there is a clear distinction of the different roles. Copying from another comment I wrote here: > These are two separate things: > * Making content available for unauthorized distribution > * Distributing unauthorized content that someone else already made available > Seeding isn't making content available, it's keeping content available.

Replied to your other comment (sorry, didn’t clock that we had two threads ongoing)

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#282
post #110

Earlier quoted context omitted.

Pray, Mr. Babbage, if you put into the machine wrong figures, will the right answers come out?

I am not able rightly to apprehend the kind of confusion of ideas that could provoke such a question.

I can. He was asking if Babbage was cheating.

You put in 2+2 - the right figures. The machine says 4 - the right answer. If you put in the wrong figures, like 3+3, will the machine still say 4? It's easy to make a machine that always says 4.

The people who asked him that question, however, probably got a different scam demonstrated to them every every. Remember the Mechanical Turk? Babbage's reply paints him very honestly. It shows that he couldn't even conceive that someone might try to trick the royal court (or whoever it was) into accepting a fake device.

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#283
post #24

The same happens with whisper-large-v3 on Chinese transcription: silence is transcribed to something like "please upvote, share and favourite this video". I suspect they trained the model on some random YouTube video without carefully picking really useful data.

oh yeah this happens a lot on reddit on videos in foreign languages

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#284
post #274

Earlier quoted context omitted.

It's just "training"

You seem to equate "training" (with scare quotes) with someone actually pirating a blu-ray, but they really aren't equivalent. Courts so far have ruled that training is fair use and it's not hard to see why. Unlike copying a movie almost verbatim (as with ripping a blu-ray), AI companies are actually producing something transformative in the form of AI models. You don't have to like AI models, or the AI companies' bu…

Who's to say why I downloaded and am now watching a movie? Is it for my enjoyment? Is it because I'm training my brain? How is me training my brain any different from companies training their LLMs?

Same goes for recording: I'm just training my skills of recording. Or maybe I'm just recording it so I can rewatch it later, for training purposes, of course.

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#285

Earlier quoted context omitted.

No, if you revolutionize both the practice and philosophy of computing and advance mankind to the next stage of its own intellectual evolution, you get to do whatever the fuck you want. Seems fair.

Hm. Not a given that it's an advance.

At the risk of stepping on a well-known land mine around here, how'd you do on the IMO problem set this year?

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#286

The fork that I've been using, WhisperX, seems to do better. I've used it on clean splits of mic tracks (ie total silence when the other is talking) with far fewer hallucinations.

WhisperX works better because it implements a robust VAD (Voice Activity Detection) preprocessing step that effectively filters out silence segments before they reach the model, preventing the hallucination triggers entirely.

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#287

Earlier quoted context omitted.

No, if you revolutionize both the practice and philosophy of computing and advance mankind to the next stage of its own intellectual evolution, you get to do whatever the fuck you want. Seems fair.

Except that the jury’s (at best) still out on whether the influence of LLMs and similarly tech on knowledge workers is actually a net good, since it might stunt our ability to critically think and problem solve while confidently spewing hallucinations at random while model alignment is unregulated, haphazard, and (again at best) more of an art than a science.

Well, if it's no big deal, you and the other copyright maximalists who have popped out of the woodwork lately have nothing to worry about, at least in the long run. Right?

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#288

Earlier quoted context omitted.

Whether it’s legal slash fair use to train on copyrighted material is only one of the questions currently being asked though. There’s a separate issue at play where these companies are pirating the material for the training process. By comparison, someone here brought up that it might be transformative fair use to write a play heavily based on Blood Meridian, but you still need to buy a copy of the book. It would sti…

If they would buy material at a large scale, the seller might require them to sign a contract that requires royalty if the material is used for training an AI. So buying legally is a way to put yourself into a trap.

What is the precedent on that kind of agreement?

The only thing I've been able to find is the note that since copyright is federal law, state contract law actually can't supersede it, to wit: if you try to put a clause in the contract that says the contract is void if I use your work to make transformative fair-use works (or I owe you a fee), that clause is functionally unenforceable (for the same reason that I don't owe you a fee if I make transformative fair-use works of your creations in general).

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#289

Earlier quoted context omitted.

I wouldn't describe it as "unusable" so much as needing to understand its constraints and how to work around them. I built a business on top of Whisper [1] and one of the early key insights was to implement a good voice activity detection (VAD) model in order to reduce Whisper's hallucinations on silence. [1] https://speechischeap.com

How does this make a profit? Whisper should be $0.006 to $0.010 per minute, but you rate less than $0.001? Do you 10x the audio?

Thanks for noticing. It took a lot of effort to optimize the pipeline every step of the way. VAD, inference server, hardware optimization, etc. But nothing that would compromise on quality. The audio is currently transcribed in its original speed. I'll be sure to publish something if I manage to speed it up without incurring any losses to the WER.

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#290
post #107
post #58

Title should be changed to "OpenAI publishes evidence they trained on pirated movies".

Of course. Piracy is legal when you have a bigger pile of money than the studios.

Isn't Piracy legal in many parts of the world?

Legally, why wouldn't they be able to do the piracy parts in one of those jurisdictions and then ship the outputs back to the mothership?

Post reply on HN