Earlier quoted context omitted.
I just tested it and it was pretty mediocre at least with my accent. I can definitely benefit from a decent app for quick note recording with a button press->transcribe->upload to gdrive/good UI app for later grepping.
Was this with the default base model, or the medium or large model? This can be specified with the —model flag.
Whisper – open source speech recognition by OpenAI
251–260 of 508 posts
Re: Whisper – open source speech recognition by OpenAI
#252Neat, https://github.com/openai/whisper - they have open-sourced it, even the model weights, so they are living up to their name in this instance. The 4 examples are stunningly good (the examples have speakers with heavy accents, speaking in foreign language, speaking with dynamic background noise, etc.), this is far and away better than anything else I've seen. Will be super curious to see other folks trying it out…
Re: Whisper – open source speech recognition by OpenAI
#253Earlier quoted context omitted.
The French version is a little contrived. The speaker is a native speaker, but the text is obviously the result of a translation from English to French, not idiomatic French. I will try to put the code to the test, see how it goes.
I'm interested in building something with this to aid my own French learning. Would love to read your findings if you end up posting it somewhere like twitter/blog!
Original:
> Mes révérends pères, mes lettres n’avaient pas accoutumé de se suivre de si près, ni d’être si étendues. Le peu de temps que j’ai eu a été cause de l’un et de l’autre. Je n’ai fait celle-ci plus longue que parce que je n’ai pas eu le loisir de la faire plus courte. La raison qui m’a obligé de me hâter vous est mieux connue qu’à moi. Vos réponses vous réussissaient mal. Vous avez bien fait de changer de méthode ; mais je ne sais si vous avez bien choisi, et si le monde ne dira pas que vous avez eu peur des bénédictins.
Transcription:
> Mes rêves errent pères, mais l'detre navais pas accoutumé de se suivre de si près ni d'detre si étendu. Le peu de temps que j'sais eu a été cause de l'de l'de l'de autre. J'sais n'detre plus longue que parce que j'sais pas eu le loisir de la faire plus courte. La raison qui m'sa obligée de me hâter vous est mieux connue qu'moi. Vos réponses vous réussissaient mal. Vous avez bien fait de changer de méthode, mais je ne sais pas si vous avez bien choisi et si le monde ne dira pas que vous avez eu peur des bénédictes.
Here there are many more mistakes, so many that the beginning of the text is unintelligible. The language from the 17th century is probably too different. Still on the "medium" model, as the large one crashes the Colab (not sure how to select a beefier machine.)
Still fascinating and exciting though.
Re: Whisper – open source speech recognition by OpenAI
#254Can this be used as a real-time transcription or is it too slow for that? Curious what anyone is using these days for a real-time transcription. It doesn't have to be perfect, but just good enough. My kids watch some youtube vidoes where people will make a mod where it converts them talking to text then look for keywords and spawn a boss in Terraria if you say the wrong keyword etc. I made a clone of that with the .N…
If your family uses Apple devices, Apple offers free on-device speech recognition. Only caveat is that it needs to be restarted every minute due to whatever stupid limitation (or bug) they've introduced. https://developer.apple.com/documentation/speech/recognizing... Also, see `requiresOnDeviceRecognition`
Re: Whisper – open source speech recognition by OpenAI
#255"[01:17.000 --> 01:32.000] Translated by Releska" when using the translate to english. That entire part of the song is instrumental. This line does not appear at all in the original transcribe only in the opus format rip.
It shows up in the yt rip in format 251 (opus), but not in format 140 (aac from youtube), nor the flac rip. All three are giving different results.
The translation quality is tied to bitrate. Same song converted to different words, the only difference being bitrates and formats. Converting my own rip with the same parameters as yt (opus @140 and then @130) didn't allow me to reproduce this error.
The model hung for a solid extra minute at the end when translating to english, the last 90ish seconds of the song took real time 60 seconds, while the entire rest took about 90. The same behavior was not observed with the transcribe.
Some of the english words are incorrect but that was expected. The first Japanese "mistake" I found was "全ては二人の" instead of "すべては ふたりの". With the left being what whisper wrote. A single random word "hey" was transcribed/translated to english even though it's the singer elongating the 園 while singing the 楽園. "落ちてゆく 二人で繋がれた二人のラグ HEY" instead of "落ちていく 鎖でつながれた 二人の楽園" .
I am using the official subtitles released on the youtube video.
It's a complex Japanese song with both japanese and english, and the original transcribe took about 20 real time seconds to start with the first line, 130 seconds for the whole song. It seems to be showing results in 20 second window increments, but this seems to depend on what it considers audio and what it is throwing away.
On my computer I wasn't able to use the large model because I ran out of VRAM, I have 8gb, not sure how much more it'd require. So I ran it with medium.
The song is False Sympathy by Mondo Grosso. The mv is suggestive, in case that matters. I grabbed a fresh audio rip from Youtube because I didn't want to take it out of my cd case.
https://www.youtube.com/watch?v=B6Y-WsgpzlQ
It is translating this version differently from the director's cut version. I ripped both as opus.
There is something weird about how it is handling the opus encoded version, as I find the same "Translated by Releska" in a wav version transcoded from the opus.
Re: Whisper – open source speech recognition by OpenAI
#256Comparing this model's word error rates to the state of the art [1] on a few common test sets: Whisper SoTA LibriSpeech test-clean 2.7% 1.8% LibriSpeech test-other 5.6% 2.9% Switchboard 13.1% 4.9% CallHome 15.8% 9.5% The authors do explicitly state that they're trying to do a lot of fancy new stuff here, like be multilingual, rather than pursuing just accuracy. [1] https://github.com/syhw/wer_are_we
I suspect Whisper is more robust than other "SOTA" models, but this release is likely leaving a fair bit of accuracy on the table considering the amount of resources OpenAI is capable of throwing at training it. Comparing the readily available test sets from the paper to some of my personal robust models (for the Talon models, this is greedy decoding, no language model): Talon Talon Talon Whisper wav2vec 2.0 28M 300M…
Re: Whisper – open source speech recognition by OpenAI
#257Japanese results looks pretty impressive! Took マッコウクジラ14頭が海岸に打ち上げられる オーストラリア(2022年9月21日) https://www.youtube.com/watch?v=bZkNIzeRBk4 Extracted audio with youtube-dl -f bestaudio https://www.youtube.com/watch\?v\=bZkNIzeRBk4 Converted into [00:00.000 --> 00:13.000] オーストラリア南部の島で、真っ向くじら14棟が海岸に打ち上げられて死んでいるのが見つかり、専門家が調査のため原地入りしました。 [00:13.000 --> 00:25.000] 原地メディアによりますと、オーストラリア南部のキング棟で、19日、少なくとも14棟の真っ向くじらが海岸に打ち上げられて死んでいるの…
Re: Whisper – open source speech recognition by OpenAI
#258Earlier quoted context omitted.
The French version is a little contrived. The speaker is a native speaker, but the text is obviously the result of a translation from English to French, not idiomatic French. I will try to put the code to the test, see how it goes.
I'm interested in building something with this to aid my own French learning. Would love to read your findings if you end up posting it somewhere like twitter/blog!
Original:
Trois mille six cents fois par heure, la Seconde
Chuchote Souviens-toi !– Rapide, avec sa voix
D'insecte, Maintenant dit Je suis Autrefois,
Et j'ai pompé ta vie avec ma trompe immonde !
Remember ! Souviens-toi ! prodigue ! Esto memor !
(Mon gosier de métal parle toutes les langues )
Les minutes, mortel folâtre, sont des gangues
Qu'il ne faut pas lâcher sans en extraire l'or !
Transcription:> Trois mille six cents fois par heure, la seconde chuchote « Souviens toi », rapide, avec sa voix d''insecte, maintenant dit « Je suis autrefois », et j''ai pompé ta vie avec ma trompe immonde. « Remember, souviens toi, prodigue, est au mémoire, mon gosier de métal, parle toutes les langues, les minutes, mortelles folâtres, sont des gangs qu''il ne faut pas lâcher sans en extraire l''or. »
Not bad! Far from perfect but it's a difficult text. Interesting that it works better with Baudelaire than Pascal.
Re: Whisper – open source speech recognition by OpenAI
#259> The source code must be the preferred form in which a programmer would modify the program. [...] Intermediate forms such as the output of a preprocessor or translator are not allowed.
If I asked a programmer from OpenAI to modify the model to better support Japanese speakers from Hokkaido, their "preferred form" of the model's source code would include the 680,000 hours of audio used to train the model.
Yes that means that there are almost no open source models and yes it's awesome that they released this and made the weights available. Just don't call it open source.
Re: Whisper – open source speech recognition by OpenAI
#260Earlier quoted context omitted.
All of your examples are limited in some way, but GPT-3 wouldn't have any meaningful limits. Stable Diffusion: Marks images as AI-generated. (invisible watermark, but still, it's there) Photoshop: Requires time & effort from a human. Fake news website: Requires time & effort from a human.
I wouldn't really say Stable Diffusion marks images as AI-generated. There's a script in the Stable Diffusion repository that will do that, but it's not connected to the model itself in a meaningful way. I use Stable Diffusion a lot and I've never touched this script. https://github.com/CompVis/stable-diffusion/blob/69ae4b35e0a...
Trivial to remove, I give you that. But AFAIK, the original repository + most forks put the watermark automatically unless you've removed it on your own.