This is really nice. I tried it out on the app and wanna start using it more. The only issue stopping me is the privacy policy. I understand why you have to retain the voice data, however, is there any way you can implement an opt-out feature? I’m just not comfortable with my voice data lingering in some servers, ready to be used for training ml models.
Launch HN: Aqua Voice (YC W24) – Voice-driven text editor
171–180 of 255 posts
Re: Launch HN: Aqua Voice (YC W24) – Voice-driven text editor
#172Since voice-to-text has gotten so good I've used it a lot more and also noticed how distracting and confusing it can be. Using Apple's dictation has a similar feel to this where you're constantly seeing something that's changing on the screen. It's kind of irritating and I don't really know what the solution is. One suggestion I have here is to have at least two different sections of the UI. One part would be the act…
Distil-whisper is incredibly fast. Realtime on a 3060 Ti, and I used it to transcribe an 11 hour audiobook in 9 minutes.
I kid. Your comment made me think of a shower thought I had recently where I wished my audiobook had subtitles.
Re: Launch HN: Aqua Voice (YC W24) – Voice-driven text editor
#173Holy crap I’m blown away by the demo. That was so easy and natural to use. I had a brush with RSI some years ago. I’m good now with an ergonomic keyboard and better habits (water, exercise) but it brings me great comfort knowing something like this exists. Thank you!
Re: Launch HN: Aqua Voice (YC W24) – Voice-driven text editor
#174This is excellent and very impressive. Have you thought about offering this as an API? I'd bet there are lots of startups that want to easily integrate better speech-to-text for conversational AI (e.g. word correction, adding punctuation, etc), and would pay you for the service. (Personally, I would! Email me at matt@syntheticdreamlabs.com if you're interested in offering it as an API, I'd be pretty curious about pri…
Re: Launch HN: Aqua Voice (YC W24) – Voice-driven text editor
#175Re: Launch HN: Aqua Voice (YC W24) – Voice-driven text editor
#176This is super impressive! I am really looking forward to the day when dictation like this can be done locally on our phones. I'd really like to do a lot of my basic messaging with my voice but the need to do corrections with that tiny keyboard means it's not much of a time saver.
Re: Launch HN: Aqua Voice (YC W24) – Voice-driven text editor
#177Re: Launch HN: Aqua Voice (YC W24) – Voice-driven text editor
#178Earlier quoted context omitted.
Awesome. Agree on the copy-paste annoyance, we're working on more clients. But I do think that the reliability needs to take a few more steps before it becomes a true keyboard replacer.
Thanks for all your hard work! Even, as a start, I found myself asking the app to copy the text to the clipboard for me without even thinking. Might be nice to be able to do that more seamlessly, just as a start? You've moved us all a lot closer to my dream: taking a long walk outside, AirPods in, and handling the day's email without even looking at a screen once.
I have a similar dream, we'll make it happen!
Re: Launch HN: Aqua Voice (YC W24) – Voice-driven text editor
#179I noticed a correction that was done retroactively in the demo
'make that H100 GPUs'
and noticed that there was only one instance of the token GPU. Hence the correction was seamless. Had there been a couple more instances I guess all GPU tokens would have been replaced by H100 GPUs. I guess you could say make that NVIDIA H100 GPUs that would be more accurate but if there were multiple instances and you needed the change only in one instance, not sure how that'd fly. I am nitpicking but this could be a common theme.
The fact this can retroactively change the text and also understand a command is quite brilliant. I don't see any trigger word for a command, so wondering if I needed the command as part of actual text, how would that work ?
Re: Launch HN: Aqua Voice (YC W24) – Voice-driven text editor
#180This is cool! Some feedback: - As others have said, "1000 tokens" doesn't mean anything to non-technical users and barely means anything to me. Just tell me how many words I can dictate! - That serif-font LaTeX error rate table is also way too boring. People want something flashy: "Up to 7x fewer errors than macOS dictation" is cool, a comparison table is not. - Similarly, ".05 Word Error Rate" has to go. Spell out w…
> People want something flashy: "Up to 7x fewer errors than macOS dictation" is cool, a comparison table is not. Respectfully disagree on this one: as a startup, you can't effectively compete with the likes of Apple on flashiness. However, the very target market of those dictating large amounts of text will include a significant number of people in academia themselves. For those people, Aqua Voice will feel relevant.…