Live data from Hacker News

FFmpeg 8.0 adds Whisper support

code.ffmpeg.org

311–320 of 341 posts

Re: FFmpeg 8.0 adds Whisper support

#311

Earlier quoted context omitted.

They do that because it increases “engagement”, not because they care about the user’s experience with the subtitles.

I did that (distracting subtitles) on one of my videos and it had a very negative response. I won't do it again, but I was puzzled because I find it much nicer than the traditional subtitle format personally. It's easier for my brain to focus on. (And no one in my test audience minded.)

Do you happen to have ADHD? That might explain the discrepancy :)

Re: FFmpeg 8.0 adds Whisper support

#312
post #154

Earlier quoted context omitted.

This is called linguist relativity (nee. The Sapir-Whorf hypothesis) and the strong form you describe has fallen out of favour in modern linguistics. A surprising number of monolingual people think their own language is the most adaptable and modern language, but this is obviously untrue. All languages evolve to fit the needs of speakers. Also, the idea that people "think in language X" is heavily disputed. One obvio…

My experience is that sometimes, for example, when I watch a lecture in a foreign language, there could be some terms for which I don't know the correct translation so I cannot think about or mention them in my native language, while I understand what they mean.

I was more focused on the experience of monolinguals (where this kind of explanation is impossible), but yes I also experience this fairly often as someone who speaks more than one language.

Re: FFmpeg 8.0 adds Whisper support

#313

May I ask, if there is a movie where English people speak English, French people speak French, and German people speak German, is there a software that can generate subtitles in English, French and German without translating anything? I mean, just record what it hears.

[dead]

Re: FFmpeg 8.0 adds Whisper support

#314
post #285

Earlier quoted context omitted.

I don’t mean to speak for OP, but it strikes me as rude to make light of someone’s disability in this way. I’d guess it has caused them a lot of frustration.

Your assumption leads you to believe that I do not also suffer from the same issue. Ever since I was in a t-bone accident and the side airbag went off right next to my head, I have a definite issue hearing voices in crowded and noisy rooms with poor sound insulation. Some rooms are much worse than others. So when I say I call it a feature, it's something I actually deal with unlike your uncharitable assumption.

Sometimes, late at night when I'm trying to sleep, and I hear the grumble of a Harley, or my neighbors staggering to their door, I wonder: why do we not have earflaps, like we do eyelids?

Re: FFmpeg 8.0 adds Whisper support

#315
post #230

Earlier quoted context omitted.

Can you give an example why it made your life that much better?

I used it like sibling commenter to get subtitles for downloaded videos. My hearing is bad. Whisper seems much better that YouTube's built-in auto-subtitles, so sometimes it is worth the extra trouble for me to download a video just to generate good subtitles and then watch it offline. I also used whisper.cpp to transcribe all my hoarded podcast episodes. Took days of my poor old CPU working at 100% on all cores (and…

This, but I want a summary about the 3 hour video first before getting spending the time on it.

Download -> generate subtitles -> feed to AI for summary works pretty well

Re: FFmpeg 8.0 adds Whisper support

#316
post #257

Earlier quoted context omitted.

> When I'm reading a transcript That's the thing though, subtitles aren't intended as full transcripts . They are intended to allow a wide variety of people to follow the content. A lot of people read slower than they would hear speech. So subtitles often need to condense or rephrase speech to keep pace with the video. The goal is usually to convey meaning clearly within the time available on screen. Not to capture e…

I regularly enable YouTube subtitles. Almost always, they are a 100% verbatim transcription, excluding errors from auto-transcription. I am not annoyed in the slightest, and in fact I very much prefer that they are verbatim. If you are too slow at reading subtitles, you can either slow down the video or train yourself to read faster. Or you can just disable the subtitles.

> If you are too slow at reading subtitles, you can either slow down the video or train yourself to read faster. Or you can just disable the subtitles.

And what are deaf people supposed to do in a cinema, or with broadcast TV?

(And I'm ignoring other uses, e.g. learning a foreign language; for that, sometimes you want the exact words, sometimes the gist, but it's highly situational; but even once you've learned the language itself, regional accents even without vocabulary changes can be tough).

Re: FFmpeg 8.0 adds Whisper support

#317
post #45

Once local transcription is in more places hopefully we can persuade content creator not to burn bouncing sub-titles into their videos. I've seen professionally produced recordings on dry and technical subjects with good sound quality where they've decided to use distracting sub-titles with no way to disable them. It seems so unnecessary if you're not making novelty videos about cats. Also local transcription allows…

Those burned in subtitles still aren’t as cool as theme-matched anime subtitles during intro music sequences from fansubs 15 years ago. Those are still cool IMO

I recently discovered that the Internet Archive has the Tomodachi fansubs of Fushigi Yugi which, at least in my experience, were the most famous example of that technique.

https://archive.org/details/tomodachi-fushigi-yugi-vhsrip

Re: FFmpeg 8.0 adds Whisper support

#318
post #293

Earlier quoted context omitted.

I don't understand how this would happen, though. It's not like it will mishear a dutch sentence as if it's english; it will correctly pick up the dutch sentence, but (since the language is auto-detected as english at the start of the segment), seemingly auto-translate that (correct and correctly heard) dutch text to english. All we need is a way to get the dutch text that's surely somewhere in there, before the tran…

Maybe try the turbo model which is transcription only. The other models were trained on x to en translations and they seem to emphasise the output language over the task token. You can get them to translate to any language even though it was never trained for that, comparatively nl-en translation is in the dataset so I'm not surprised it's doing that.

Hey, good tip, thanks a lot!

Re: FFmpeg 8.0 adds Whisper support

#319
post #62

Earlier quoted context omitted.

Isn't that a bit much for ASR models? Humans can't handle simultaneous multilingual dictation task either, I have to stop and reinitialize ears before switching languages between English and my primary one.

In South Asia, it's quite common for people to speak a combination of their local language and English. Not just alternating sentences between the two languages, but in fact, constructing sentences using compound phrases from the two languages. "Madam, please believe me, maine homework kiya ha" [I did my homework].

This is common in the southwestern part of the US too. My partner and her friends she grew up with will have conversations that fluidly pick phrases and vocab from either Spanish or English depending on what words happen to be the easiest to pull from their brain. It's wild to listen to.

Re: FFmpeg 8.0 adds Whisper support

#320
post #257

Earlier quoted context omitted.

I regularly enable YouTube subtitles. Almost always, they are a 100% verbatim transcription, excluding errors from auto-transcription. I am not annoyed in the slightest, and in fact I very much prefer that they are verbatim. If you are too slow at reading subtitles, you can either slow down the video or train yourself to read faster. Or you can just disable the subtitles.

> If you are too slow at reading subtitles, you can either slow down the video or train yourself to read faster. Or you can just disable the subtitles. That's just plain tone deaf, plain and simple. I was not talking about myself, or just youtube. You are not everyone else, your use case is not everyone else their use case. It really isn't that difficult.

You made a bet and lost. Things are difficult.
Post reply on HN