Is there even a good offline version of this? There are some opensource tools for speed-to-text but what about batch processing of audio files?
Here's an example using GNU parallel: http://voice2json.org/recipes.html#parallel-wav-recognition
11–20 of 21 posts
Is there even a good offline version of this? There are some opensource tools for speed-to-text but what about batch processing of audio files?
Here's an example using GNU parallel: http://voice2json.org/recipes.html#parallel-wav-recognition
What’s the pricing? What speech-to-text engine is being used? Clicking on the Sign Up button on iOS Safari does nothing. Clicking on the Get Started button takes me to an Upload Video form - not what I expected from a mp3-to-text service.
Is there even a good offline version of this? There are some opensource tools for speed-to-text but what about batch processing of audio files?
You may be interested in voice2json for offline batch processing: https://voice2json.org Here's an example using GNU parallel: http://voice2json.org/recipes.html#parallel-wav-recognition
"MP3 to Text" seems very inaccurate since you can only upload video files. In fact uploading an .mp3 file shows "File type not supported". edit: I get it, OP just keeps submitting his service with different descriptions until one gets some upvotes. Only took 25 tries to get 30 points. Shameful.
Is there even a good offline version of this? There are some opensource tools for speed-to-text but what about batch processing of audio files?
You may be interested in voice2json for offline batch processing: https://voice2json.org Here's an example using GNU parallel: http://voice2json.org/recipes.html#parallel-wav-recognition
> Sets of voice commands that are described well by a grammar
> Commands with uncommon words or pronunciations
> Commands or intents that can vary at runtime
Doesn't sound like what you'd want for a generic transcription service.
Earlier quoted context omitted.
You may be interested in voice2json for offline batch processing: https://voice2json.org Here's an example using GNU parallel: http://voice2json.org/recipes.html#parallel-wav-recognition
> voice2json is optimized for: > Sets of voice commands that are described well by a grammar > Commands with uncommon words or pronunciations > Commands or intents that can vary at runtime Doesn't sound like what you'd want for a generic transcription service.
Users have reported good accuracy with the English Deepspeech profile: https://github.com/synesthesiam/voice2json-profiles
MP3 to text? Why does it ask me to upload a video?
Very cool but how do I know what languages supported? It says "VEED is able to recognise and transcribe languages from all over the world - English, Spanish, French, Chinese, and many more". From my experience with NLP/AST the tricky part is models for some less common languages.
"MP3 to Text" seems very inaccurate since you can only upload video files. In fact uploading an .mp3 file shows "File type not supported". edit: I get it, OP just keeps submitting his service with different descriptions until one gets some upvotes. Only took 25 tries to get 30 points. Shameful.