Live data from Hacker News

Whisper – open source speech recognition by OpenAI

openai.com

331–340 of 508 posts

Re: Whisper – open source speech recognition by OpenAI

#331

Earlier quoted context omitted.

Did you try translating them to english? I want to see if you get a similar error as me with a random phrase "Translated by Releska" showing up.

It's called hallucination. As the model is trained on unsupervised data, such errors do seldom happen. The model picks up that such phrases occur in translations and inserts them even if they do not appear in the source. This is described in the paper.

I came across it during a silent/instrumental portion in the song I was testing. I asked only because I am curious how frequently the error might show up, I don't expect it to be very common. It's looking at phrase level instead of word level timestamps which is going to make it hard to tokenize music. I asked simply because the parent comment also tested on Japanese.

Re: Whisper – open source speech recognition by OpenAI

#332
Got my hopes high that there's finally an open source solution that can deal with Georgian language, only to get my hopes brutally destroyed. It successfully detects a language and then produces garbage. Passing language manually produced similar results.

Result of my own recording:

  Detected language: georgian
   ᔨᴉᴉ�ちゃんᓁᔇ � remnants ᡔ� founding ហ�ockey� slee សᕁ �eling ភᕩ�icularly អᕖᕤ�APPLAUSEPS ថ�Dav頻道 ប�DING� Możai បፘ្ទក ុក ឵� orchestral ុក ឵� arter ូ� Brettំ � 
  hilarious ល ឬ ᔼ� vårក បក ្៙ � Poll statements ឭ᪨្pson. ჩჩრუესი჏მეისლემვეერრშუეაირელმირისასასსსესსერერსივეესრრილმეხრე რეიმიმეფემსესე�
Results of clear Georgian audio [1].

On tiny model:

  Detected language: georgian
  [00:00.000 --> 00:21.560]  én
  [00:21.560 --> 00:23.240] 我伦伦…
  [00:23.280 --> 00:43.720] 我伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦伦因为b forestry

On medium model:

  Detected language: georgian
   სრჱირესრრრრრრრრრრრრრნსსსრრრრრეე რრირრრრრრრრრე რსრნგნრრრრსრრრრრრრორრრრრრრრრრრ� ḵḸḇḤḾḤḾḤḾḤḾḤḾḤḾḤḾḤḾḾḤḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾ� ḥḾḼḥḾ 
  ḥḾḾ ḥḾḾ ḤḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾḾ� ḲḵḽḻḽḾ Ḫḵḽḻḽ so� ḻḽḽ ḻḽḻḻḽ ḱᴇ᷻ᵒ ḳᶟᄤḱ ḯᵁ Ḳᴄᴍᴆ Ḧᴍ� Ḧᵒ ḳᴍᴇ ḽᴄᴍᴛᴄ Ḧᴇᴆ ḳᵗᴇ ḽḮᴆ Ḫᴇᴾ ḿᴏᴇᴄᴄᴏ 
  ច�izar� wait �ห� examined ᑇទមះៈេំ supervision ង� იეეეეეეეეეეეეეეეეე მაეე ეაეეეეეეეეეეეეეეეეეეეე დაეეეეეეეეეეეეე უეეეეეეეეეეეეე ეა� ჆ მიი სმეიი მმიეი Ⴢქ სიიეი 
  სავიე სიიითთიიმემი, რაეე სიიმე სიიი ღიიიიწეირი საეიეიი სიიეი სი� ვეეფვეიიიე ქლეეშეეროეეეეეეეეეეეეე. ეგეზ ეყაკშეიეეეეეეეეეეეეეეეეეეეეეეეეეეეეეა, ნრროპიროო მმუმინ 
  სეეკნფეე სეეჍიგოშ სჟებიმელელეეკირპიე სემეიმე სეეიმმმ სეენემეეი სე� ᑦ� Famose m인데요 hqe bywall jaini threshold ji jani den poder vlogging bywall Take the text Ba 
  tou yodamj je te shake ba te shake baou contour but whatever Baou cube baou cup Baou rope Baou people Qeful Qeful იმიიიმიბთმითიიითიიიიიიიი 
  რაოეოოოენპეეეიეიიიიიიიიიომიიიიიიიიი რიიიიიიიიიიიმიი� ნსეეეეეეეეეეეეეეე სარეეეეეეეეეეეეეეეეეეეეეეეეეეეეეეეეეეეეეეეეეეეეეეეეეეეეეე� მጇი჏ვ ეეეიდჼვვ ნაბდადებ 
  ლმირეეეეფედუივევეეეიიეეეეე რარეიეეეევეეეეევეე სარრეეეეეეეეეეეეეეეეეეეეეეეეეეე ხშიიიიიიიიიიიიი ლიიიიიიი ლიიიიიიიიიი ლიიი ლიიიიიიი ლაიიიიი ეიიიიიიიიიიიიიიი იიიი მ�

I've also tested it on few other audio inputs and it failed to produce meaningful results on all of them with all models.

There was one case with another audio [2] and tiny model, where it got at least some words close to their phonetic values, but printed them in cyrillic instead of Georgian and tried to interpret some Georgian words as Russian:

  whisper audio.wav --language Georgian --task transcribe --model tiny
  [00:00.000 --> 00:02.000]  «Зураб Герча Джапарзис Ганц Хатеваром
  [00:02.000 --> 00:04.000]  умерен цупасу Хизгеблоту кащепаста
  [00:04.000 --> 00:06.000]  а опозационермии член шонахлари
  [00:06.000 --> 00:07.000]  с дрородисат Сакартолом
  [00:07.000 --> 00:09.000]  с акутаритеритория бюнда дай бронос
  [00:09.000 --> 00:10.000]  та тасовый торуси сам кадр
  [00:10.000 --> 00:12.000]  Сакартоломший ровно украйенисту
  [00:12.000 --> 00:13.000]  щойго екнебо
  [00:13.000 --> 00:14.000]  амсясахеб кирчи метитаусу
  [00:14.000 --> 00:15.000]  хлебислидерма
  [00:15.000 --> 00:17.000]  уцноктангадацема щейсяа уградунца
  ...



[1] https://www.youtube.com/watch?v=rE_zx_6RhL0 [2] https://www.youtube.com/watch?v=elrXgO8hjtI

Re: Whisper – open source speech recognition by OpenAI

#333

Earlier quoted context omitted.

For real. The way people normally speak, with backtracking, repetition, restarting sentences, or stopping mid sentence and starting a new one with entirely different nouns or entire subjects is perfectly normal in synchronous conversation and isn't jarring, but written down as is, it's like 40% noise.

For a good example of this, read ANY of trumps speaches transcribed.

I mean if you want to make it unnecessarily political, Biden's are worse: https://www.youtube.com/watch?v=3bWM1zsnTJc

Re: Whisper – open source speech recognition by OpenAI

#336

Earlier quoted context omitted.

Ran a few other songs through it and found one obvious mistranscription: "He's the bedroom cosmic rocker" (should be "He's the veteran cosmic rocker" in Veteran Cosmic Rocker by The Moody Blues) I also noticed that it's a little on the conservative side for detecting speech; all songs were missing at least part of one line.

Ran it on Juicy by The Notorious B.I.G and results were considerably worse than my mix of prog-rock and british invasion music I had tried before, though at least some of that is due to the number of proper-nouns in that song. It took about 1000 CPU-minutes for this 5 minute song on my Ryzen 2700 with 12 OpenMP threads (about 100 minutes wall-clock).

Here's the output of

    whisper never-gonna-give-you-up.mp3 --language English --model small

    [00:00.000 --> 00:27.000]  We're no strangers to love You know the rules and so do I
    [00:27.000 --> 00:35.000]  I feel commitments while I'm thinking of You wouldn't get this from any other guy
    [00:35.000 --> 00:43.000]  I just wanna tell you how I'm feeling Gotta make you understand
    [00:43.000 --> 00:47.000]  Never gonna give you up Never gonna let you down
    [00:47.000 --> 00:53.000]  Never gonna run around and desert you Never gonna make you cry
    [01:00.000 --> 01:09.000]  We've known each other for so long Your heart's been aching but you're too shy to say
    [01:09.000 --> 01:17.000]  Inside we both know what's been going on We know the game and we're gonna play it
It was running for quite a long time (20 minutes) on my admittedly low-budget specs.

Note that I did not omit 00:53.000 -> 01:00.000.

Shouldn't there be some type of unintelligible warning since it wasn't able to transcribe that part?

Re: Whisper – open source speech recognition by OpenAI

#337

A notebook is available to try with your microphone on Colab here: https://colab.research.google.com/drive/1nBZ-pDIaIi3N1DIIXvJ... I'm surprised by the quality on non-English languages, given that 80+% of the training data is English, and the rest is split between tens of languages.

How do you get this to translate instead of just transcribe?

Re: Whisper – open source speech recognition by OpenAI

#338

That example at the top of the page (speed talking) blew me away. He started talking, I was stunned for a minute, then realised yes, it really was English, and I just burst out laughing. That's so, so far beyond the previous state-of-the-art, it's absurd.

It's a micromachines ad from the '80s. He talked like that in all of them! As for speed, to a computer we don't talk very fast, not even that guy. I wonder if it could handle Rap God by Eminem....Let's find out!

Did you find out :D?

Re: Whisper – open source speech recognition by OpenAI

#339
post #34

I just tested the model [1] using an RTX3090, trying to translate a french text I found here [2]. Some observations: - The full translation of the 6:22 minute video takes about 22 seconds (17x real time) - It recognizes the language by default (and did a good job to recognize it was french audio) - MIT License [3]! - The quality of the transcription is good, but not perfect. - The quality of the translation (if you d…

How did you get it to use the GPU? I have it running right now and it's not touching the GPU.

--device "cuda"

Re: Whisper – open source speech recognition by OpenAI

#340

Earlier quoted context omitted.

How did you get it to use the GPU? I have it running right now and it's not touching the GPU.

--device "cuda"

My version of pytorch didn't have CUDA. I had to install conda to get it, and now it's currently installing.

Whatever the default version that `pip install git+https://github.com/openai/whisper.git` grabbed didn't include it by default.

Post reply on HN