Earlier quoted context omitted.
That's fine for training your own model, but I don't think you could distribute the training set. That seems like a clear copyright violation, against one of the groups that cares most about copyright. Maybe you could convince a couple of indie creators or state-run programs to licence their audio? But I'm not sure if negotiating that is more efficient than just recording a bit more audio, or promoting the project to…
That's fine for training your own model, but I don't think you could distribute the training set. That seems like a clear copyright violation, against one of the groups that cares most about copyright. I'm not sure that is a clear copyright violation. Sure, at a glance it seems like a derivative work, but it may be altered enough that it is not. I believe that collages, and reference guides like cliff notes are both…
Voice2json: Offline speech and intent recognition on Linux
61–70 of 114 posts
Re: Voice2json: Offline speech and intent recognition on Linux
#62Earlier quoted context omitted.
I wonder if they use movies and tv; recordings where the script is already available.
That's fine for training your own model, but I don't think you could distribute the training set. That seems like a clear copyright violation, against one of the groups that cares most about copyright. Maybe you could convince a couple of indie creators or state-run programs to licence their audio? But I'm not sure if negotiating that is more efficient than just recording a bit more audio, or promoting the project to…
Re: Voice2json: Offline speech and intent recognition on Linux
#63I wonder if it would be possible to map vim keybindings to sounds and effectively drive the editor with the mouth when the hands are otherwise occupied. It might be possible to use sounds that compose into pronounceable words with minimal syllables for combinations. What would vim bindings look like as a concise command language suited to human vocalization? E.g. maybe "dine" maps to d$ and "chine" to c$. So as in ke…
Re: Voice2json: Offline speech and intent recognition on Linux
#64Earlier quoted context omitted.
I wonder if they use movies and tv; recordings where the script is already available.
I expect that wouldn't be perfect, though. Sometimes the cut that makes it into the final product doesn't exactly match the script. Sometimes it's due to an edit, other times it's due to an actor saying something similar to but not exactly what the script says, but the director deciding to just go with it. What might work better is using closed captions or subtitles, but I've also seen enough cases where those don't…
Re: Voice2json: Offline speech and intent recognition on Linux
#65Earlier quoted context omitted.
GP is not talking about the model but about the training data set.
I am aware, I'm asking if the model, however, is infringing. Surely you can't distribute them in a dataset but is training on copyrighted data legal, and can you distribute that model?
Re: Voice2json: Offline speech and intent recognition on Linux
#66Good FLOSS speech recognition and TTS is badly needed. Such interaction should not be left to an oligoply with bad history of not respecting users freedoms and privacy.
Mozilla CommonVoice is definitely trying. I always do a few validations and a few clips if I have a few minutes to spare, and I recommend everyone does. They need volunteers to validate and upload speech clips to create a dataset. https://commonvoice.mozilla.org/en
Re: Voice2json: Offline speech and intent recognition on Linux
#67Earlier quoted context omitted.
I like the idea, and decided to try doing some validation. The first thing I noticed is that it asks me to make a yes-or-no judgment of whether the sentence was spoken "accurately", but nowhere on the site is it explained what "accurate" means, or how strict I should be. (The first clip I got was spoken more or less correctly, but a couple of words are slurred together and the prosody is awkward. Without having a goo…
After listening to about 10 clips your point becomes abundantly clear. One speaker, who sounded like they were from the mid-west United States, was dropping the S off words in a couple clips. I wasn't sure if it was misreads or some accent I'd never heard. Another speaker, with a thick accent that sounded European, sounded out all the vowels in circuit. Had I not had the line being read, I don't think I'd have unders…
[¹] If that's even fair given it's a dialect in its own right - Americans also say things differently than I would as a 'Britisher'
Re: Voice2json: Offline speech and intent recognition on Linux
#68Good FLOSS speech recognition and TTS is badly needed. Such interaction should not be left to an oligoply with bad history of not respecting users freedoms and privacy.
Re: Voice2json: Offline speech and intent recognition on Linux
#69Earlier quoted context omitted.
Except they have 12k hours of audio, when really they could do with 12B hours of audio...
Then you need a lot of people that listen to those 12B hours of audio, and multiple listeners agree for each chunk of audio that what is spoken corresponds to the transcript.
Re: Voice2json: Offline speech and intent recognition on Linux
#70I wonder if it would be possible to map vim keybindings to sounds and effectively drive the editor with the mouth when the hands are otherwise occupied. It might be possible to use sounds that compose into pronounceable words with minimal syllables for combinations. What would vim bindings look like as a concise command language suited to human vocalization? E.g. maybe "dine" maps to d$ and "chine" to c$. So as in ke…
https://youtu.be/8SkdfdXWYaI?t=600 this guy is already there: Slurp slap scratch buff yank