Mozilla Overhauls Speech-To-Text Contribution Interface
voice.mozilla.org
Mozilla Overhauls Speech-To-Text Contribution Interface
1–10 of 46 posts
Re: Mozilla Overhauls Speech-To-Text Contribution Interface
#2Re: Mozilla Overhauls Speech-To-Text Contribution Interface
#3Re: Mozilla Overhauls Speech-To-Text Contribution Interface
#4My biggest issue with this project is that more than half of the contributions are from people who have failed to record correctly, or who are not fluent in English.
Re: Mozilla Overhauls Speech-To-Text Contribution Interface
#5My biggest issue with this project is that more than half of the contributions are from people who have failed to record correctly, or who are not fluent in English.
The English (in)fluency is more of a feature, though, than a bug. The goal isn't to produce a speech-to-text system that can recognize a perfectly miked BBC announcer. It's to be able to recognize a wide variety of people speaking fairly naturally in imperfect conditions, using whatever accent they use for casual speech.
Re: Mozilla Overhauls Speech-To-Text Contribution Interface
#6Re: Mozilla Overhauls Speech-To-Text Contribution Interface
#7It's awesome that the dataset is offered with a CC-0 license: https://voice.mozilla.org/en/data, does anyone know if it includes the answers from the survey? I have a limited bandwidth internet, so I haven't checked it out yet. In particular, I'm wondering if there's user data on whether they picked yes or no, to implement "troll detection" - people who click one option all the time.
Trolling of crowdsourced data isn't unheard of. (NSFW language: https://www.reddit.com/r/pics/comments/cygfx/4chan_is_using_...)
Re: Mozilla Overhauls Speech-To-Text Contribution Interface
#8My biggest issue with this project is that more than half of the contributions are from people who have failed to record correctly, or who are not fluent in English.
That's what the second pass is for, right, to screen out actually unintelligible or misrecorded entries. The English (in)fluency is more of a feature, though, than a bug. The goal isn't to produce a speech-to-text system that can recognize a perfectly miked BBC announcer. It's to be able to recognize a wide variety of people speaking fairly naturally in imperfect conditions, using whatever accent they use for casual…
Wait what? The headline is about text-to-speech aka speech synthesis, not speech recognition (speech-to-text.) Are they trying to do both? It seems to me that you'd train both using different sorts of datasets. If you wanted TTS to be intelligible to the most number of people, training to to speak like a 'perfectly miked BBC announcer' is probably exactly what you'd want to do.
Train it to recognize many regional accents, but train it to speak with the most prevalent and universally understood accent you can find. So either BBC English or Californian/Hollywood English.
Although traditionally TTS engines have shipped with numerous voices, such that you can select either a British or an America accent for the English voice. It may be worthwhile to have other English accents too, maybe one for India (125 million speakers.) But if you trained a TTS engine to have a computer amalgamation of all possible English accents I really doubt the result will be considered high quality by anybody.
Re: Mozilla Overhauls Speech-To-Text Contribution Interface
#9Earlier quoted context omitted.
That's what the second pass is for, right, to screen out actually unintelligible or misrecorded entries. The English (in)fluency is more of a feature, though, than a bug. The goal isn't to produce a speech-to-text system that can recognize a perfectly miked BBC announcer. It's to be able to recognize a wide variety of people speaking fairly naturally in imperfect conditions, using whatever accent they use for casual…
> "The goal isn't to produce a speech-to-text system that can recognize a perfectly miked BBC announcer." Wait what? The headline is about text-to-speech aka speech synthesis, not speech recognition (speech-to-text.) Are they trying to do both? It seems to me that you'd train both using different sorts of datasets. If you wanted TTS to be intelligible to the most number of people, training to to speak like a 'perfect…
Re: Mozilla Overhauls Speech-To-Text Contribution Interface
#10Earlier quoted context omitted.
That's what the second pass is for, right, to screen out actually unintelligible or misrecorded entries. The English (in)fluency is more of a feature, though, than a bug. The goal isn't to produce a speech-to-text system that can recognize a perfectly miked BBC announcer. It's to be able to recognize a wide variety of people speaking fairly naturally in imperfect conditions, using whatever accent they use for casual…
> "The goal isn't to produce a speech-to-text system that can recognize a perfectly miked BBC announcer." Wait what? The headline is about text-to-speech aka speech synthesis, not speech recognition (speech-to-text.) Are they trying to do both? It seems to me that you'd train both using different sorts of datasets. If you wanted TTS to be intelligible to the most number of people, training to to speak like a 'perfect…
For what it's worth, they do ask you to create a profile after your fifth sample, and that profile includes an "accent" section.