Live data from Hacker News

Whisper – open source speech recognition by OpenAI

openai.com

351–360 of 508 posts

Re: Whisper – open source speech recognition by OpenAI

#351

This is awesome. But I really want the other way. To be able to give it text and hear the speech. A TTS (text to speech). As a language learner, the ability to create my own sentences (based on existing ones I have, in changing a word here or there). Would be amazing. How long till we have this I wonder. I know I could use a service to do this currently. But having something running locally, I'd prefer. Hopefully som…

Likewise, TTS is what I really want. My goal is to be able to create audio books from text. I've been using Amazon Polly and it's acceptable quality, but I would be ecstatic to be able to do it locally on my own hardware.

Re: Whisper – open source speech recognition by OpenAI

#352

Earlier quoted context omitted.

Kaldi is an open, pluggable framework and is a ton more flexible and powerful than this. It's used by hundreds of teams, including a number of consumer tech companies you've heard of. They're not going to move to this over it. Especially because ASR is a living organism. You have to constantly update your language model as new people, ideas, and words move into the normal lexicon. As people start talking about "COVID…

Kaldi just is not fast or high quality enough compared to other modern alternatives like wav2letter. I appreciate that it is more flexible than this, it certainly is - but I am not so sure about "powerful."

[deleted]

Re: Whisper – open source speech recognition by OpenAI

#353

Earlier quoted context omitted.

Ran it on Juicy by The Notorious B.I.G and results were considerably worse than my mix of prog-rock and british invasion music I had tried before, though at least some of that is due to the number of proper-nouns in that song. It took about 1000 CPU-minutes for this 5 minute song on my Ryzen 2700 with 12 OpenMP threads (about 100 minutes wall-clock).

Here's the output of whisper never-gonna-give-you-up.mp3 --language English --model small [00:00.000 --> 00:27.000] We're no strangers to love You know the rules and so do I [00:27.000 --> 00:35.000] I feel commitments while I'm thinking of You wouldn't get this from any other guy [00:35.000 --> 00:43.000] I just wanna tell you how I'm feeling Gotta make you understand [00:43.000 --> 00:47.000] Never gonna give you u…

Model small is about as good at recognizing lyrics as an untrained Newton was at recognizing handwriting.

Here's a comparison of Basket Case by Greenday:

Small:

    [00:00.000 --> 00:05.000]  Do you have the time to listen to me whine
    [00:05.000 --> 00:10.000]  About nothing and everything I'll have once?
    [00:11.000 --> 00:16.000]  I am one of those melodramatic fools
    [00:16.000 --> 00:20.000]  Neurotic to the bone, no doubt about it
    [00:23.000 --> 00:27.000]  Sometimes I give myself the creeps
    [00:27.000 --> 00:32.000]  Sometimes my mind plays tricks on me
    [00:32.000 --> 00:38.000]  It all keeps headed up, I think I'm pregnant
    [00:38.000 --> 00:43.000]  And I'm just paranoid, I'm just stuck
    [00:47.000 --> 00:52.000]  I went to a shrink to have a life like my dreams
    [00:52.000 --> 00:57.000]  She says it's like a sex that's bringing me down
    [00:57.000 --> 01:03.000]  I went to a whore, he said my life's a bore
    [01:03.000 --> 01:08.000]  Choked with my widest buzz that's bringing her down
    [01:10.000 --> 01:14.000]  Sometimes I give myself the creeps
    [01:15.000 --> 01:19.000]  Sometimes my mind plays tricks on me
    [01:19.000 --> 01:25.000]  It all keeps headed up, I think I'm pregnant
    [01:25.000 --> 01:30.000]  And I'm just paranoid, I'm just stuck
    [01:30.000 --> 01:48.000]  Grasping to control, it's all I better hold on
    [02:08.000 --> 02:12.000]  Sometimes I give myself the creeps
    [02:13.000 --> 02:17.000]  Sometimes my mind plays tricks on me
    [02:18.000 --> 02:23.000]  It all keeps headed up, I think I'm pregnant
    [02:23.000 --> 02:30.000]  And I'm just paranoid, I'm just stuck
    [02:53.000 --> 03:13.000]  Thanks for watching!

Medium:

    [00:00.000 --> 00:05.000]  Do you have the time to listen to me whine
    [00:05.000 --> 00:10.000]  About nothing and everything all at once?
    [00:11.000 --> 00:16.000]  I am one of those melodramatic fools
    [00:16.000 --> 00:20.000]  Neurotic to the bone, no doubt about it
    [00:23.000 --> 00:27.000]  Sometimes I give myself the creeps
    [00:27.000 --> 00:32.000]  Sometimes my mind plays tricks on me
    [00:33.000 --> 00:36.000]  It all keeps adding up
    [00:36.000 --> 00:39.000]  I think I'm cracking up
    [00:39.000 --> 00:41.000]  Am I just paranoid?
    [00:41.000 --> 00:43.000]  Am I just sad?
    [00:47.000 --> 00:50.000]  I went to a shrink
    [00:50.000 --> 00:53.000]  To analyze my dreams
    [00:53.000 --> 00:58.000]  She says it's lack of sex that's bringing me down
    [00:58.000 --> 01:01.000]  I went to a whore
    [01:01.000 --> 01:04.000]  He said my life's a bore
    [01:04.000 --> 01:09.000]  So quit my whining cause it's bringing her down
    [01:10.000 --> 01:14.000]  Sometimes I give myself the creeps
    [01:16.000 --> 01:20.000]  Sometimes my mind plays tricks on me
    [01:20.000 --> 01:23.000]  It all keeps adding up
    [01:23.000 --> 01:26.000]  I think I'm cracking up
    [01:26.000 --> 01:28.000]  Am I just paranoid?
    [01:28.000 --> 01:30.000]  Am I just sad?
    [01:40.000 --> 01:44.000]  Grasping to control
    [01:44.000 --> 01:50.000]  So I better hold on
    [02:07.000 --> 02:11.000]  Sometimes I give myself the creeps
    [02:11.000 --> 02:16.000]  Sometimes my mind plays tricks on me
    [02:16.000 --> 02:19.000]  It all keeps adding up
    [02:19.000 --> 02:22.000]  I think I'm cracking up
    [02:22.000 --> 02:24.000]  Am I just paranoid?
    [02:24.000 --> 02:52.000]  Am I just sad?
    [02:54.000 --> 02:58.000]  Thanks for watching!

Re: Whisper – open source speech recognition by OpenAI

#354
post #5

Neat, https://github.com/openai/whisper - they have open-sourced it, even the model weights, so they are living up to their name in this instance. The 4 examples are stunningly good (the examples have speakers with heavy accents, speaking in foreign language, speaking with dynamic background noise, etc.), this is far and away better than anything else I've seen. Will be super curious to see other folks trying it out…

It seems far from good with mixed language content, especially with English and Japanese together. The timestamps are far from perfect. It's far from perfect. It's nowhere close to human for the more ambiguous translations that depend on context of word. It's far below what anyone that spoke either language would consider acceptable. Maybe it's unfair to use music, but music is the most realistic test of whether it's…

Some music is hard for even people to make out the lyrics to.

Re: Whisper – open source speech recognition by OpenAI

#355
post #5

Neat, https://github.com/openai/whisper - they have open-sourced it, even the model weights, so they are living up to their name in this instance. The 4 examples are stunningly good (the examples have speakers with heavy accents, speaking in foreign language, speaking with dynamic background noise, etc.), this is far and away better than anything else I've seen. Will be super curious to see other folks trying it out…

Can't wait to see twelve new $49.99/mo speech parser services pop up in the next few weeks.

Make hay before Google gives away free hay.

That said there is value in integration of this into other things.

Re: Whisper – open source speech recognition by OpenAI

#358
post #123

Like every model I've seen there is something like this: >>A decoder is trained to predict the corresponding text... Prediction of expected text in the context of the previous text. While this is valuable in casual transcription, it can be extremely dangerous in serious contexts. From personal experience, having given a deposition with an "AI" transcription, it will literally reverse the meanings of sentences. This i…

Do you have a demo audio clip for this? I'd be interested to see how it looks in practice.

Sorry, I don't have anything available.

One item I remember was that I said "Dr Kemeny" in relation to Dartmouth College (he was a famous mathematician, invented the BASIC programming language and was president of the college). It replaced those instances with "Jack Kennedy".

In another instance, I said that "Evidently, you have a reading comprehension problem.". It replaced it with "Evidently, I have a ...", completely reversing the meaning.

There was zero problems with the microphones or audio, and it was not rushed or mumbled talk. There were 80+ other examples over a few hours of talking, and some from other speakers. And those were just the obvious ones I could catch.

Another massive problem with this technology is that a human stenographer can notice when s/he missed something and didn't hear and ask the speaker to repeat or clarify what was said, and will often during a pause request clarification on spelling of names, addresses, etc. In contrast, this "AI" technology just barges ahead ASSuming that it knows what it is doing and inserts literally whatever sounds good in the transcript, completely silent that it doesn't have a clue.

Having seen this up close, I'm of the strong opinion that anyone foisting this software on the market without huge warnings that this is not usable for any critical functions is, basically a fraud. They know or certainly should know that these failures not only exist but are common and systemic, yet they barge along like it is OK. It is not.

Re: Whisper – open source speech recognition by OpenAI

#360
post #5

Neat, https://github.com/openai/whisper - they have open-sourced it, even the model weights, so they are living up to their name in this instance. The 4 examples are stunningly good (the examples have speakers with heavy accents, speaking in foreign language, speaking with dynamic background noise, etc.), this is far and away better than anything else I've seen. Will be super curious to see other folks trying it out…

It was already better. I edit a podcast and have > a decade of pro audio editing experience in the film industry, and I was already using a commercial AI transcription service to render the content to text and sometimes edit it as such (outputting edited audio). Existing (and affordable) offerings are so good that they can cope with shitty recordings off a phone speaker and maintain ~97% accuracy over hour-long conve…

I'm not sure if you've tried Descript, but their ML-based "Studio Sound" filter makes bad audio sound like it was recorded and edited nicely.
Post reply on HN