Live data from Hacker News

Project Common Voice

voice.mozilla.org

11–20 of 61 posts

Re: Project Common Voice

#12
post #9

This looks great! I use voice control to program on occasion due to an rsi injury. The standard stack for this is a mess due to closed source systems that aren't designed for voice programmers. A good open solution could really save me from a lot of headaches.

Previously, there was VoxForge [0], but it seems dead. At least, I failed to contribute my voice there. Mozilla getting into this space is good news indeed.

[0] http://www.voxforge.org/home/read

Re: Project Common Voice

#13
post #8

Is the data going to be freely available as well? It's a little unclear whether they intend to make it separately available or not.

I agree that this should be documented better but looking at their terms of service, the recordings submitted look to be under CC0 [1].

The relevant blurb:

    Your Contributions and Release of Rights
    
    By submitting your recordings, you waive all copyrights and
    related rights that you may have in them, and you agree to
    release the recordings to the public under CC-0. This means
    that you agree to waive all rights to the recordings
    worldwide under copyright and database law, including moral
    and publicity rights and all related and neighboring rights.

[1] https://voice.mozilla.org/terms

Re: Project Common Voice

#16
post #2

And... 503'd. I didn't catch what the intended use case was before it died, but I'm guessing computer generated voice? Most of the computer generated stuff I've seen uses trained actors. Which neatly avoids the problem of trying to reconcile a myriad of accents and dialects, which was immediately apparent from the first two samples I tried. edit: back up, seems to be about voice recognition, which this could help wit…

Actually, based on the site content I think they're using it to create an archive of speech data to train speech recognition systems.

That is correct. The DeepSpeech project, https://github.com/mozilla/DeepSpeech will use this data to train and validate open source / freely available speech to text models. The training data, along with the trained models will be made available for free to all users and researchers alike.

Re: Project Common Voice

#18
post #7

Cool project, really aligned with the mission of Mozilla, and with a pleasant UX. And if you're a non-english speaker like me validating sentences is a nice way of improving your comprehension.

Yep! Though us non-native speakers should really be recording, too. So we're not left behind in voice recognition.

Re: Project Common Voice

#19
post #17

Any idea why the duplicate detection did not work for this link: https://news.ycombinator.com/item?id=14786881 Anyhow: these should be merged (even though there is no discussion on the other submission)

I think that's why. When the previous thread is not very active, dup's are allowed. Not sure though.

Re: Project Common Voice

#20
I hope this data will be used purely for voice recognition purposes and not for voice generation, or we'll be stuck with robots talking with this horrible gurgling and clicking accent due to poor recording conditions of most participants!
Post reply on HN