It's amazing how good open-weight STT and TTS have gotten, so there's no need to pay for Wispr Flow, Superwhisper, Eleven-Labs etc. Sharing my setup in case it may be useful for others; it's especially useful when working with CLI agents like Code Code or Codex-CLI: STT: Hex [1] (open-source), with Parakeet V3 - stunningly fast, near-instant transcription. The slight accuracy drop relative to bigger models is immater…
Is Hex MacOS only?
Audio is the one area small labs are winning
91–100 of 109 posts
Re: Audio is the one area small labs are winning
#92Re: Audio is the one area small labs are winning
#93While on a walk with mobile phone + earphones, dump an article/paper/HN-Post/github-repo into the mobile chat app (chat-gpt, claude or gemini), and use voice mode to have it walk you through it conversationally, so you can ask follow up questions during the walk-thru and the AI would do web-search etc. I know I could do something like this with NotebookLM, but I want to engage in the conversation, and NotebookLM does have interactive mode but it has been super-flaky to say the least.
I pay for ChatGPT Pro and the voice mode is really bad: it pretends to do web searches and makes up things, and when pushed says it didn't actually read the article. Also the voice sounds super-condescending.
Gemini Pro mobile app - similarly refuses to open links and sounds as if it's talking to a baby.
Claude mobile app was the best among these - the voice is very tolerable in terms of tone, but like the others it can't open links. I does do web searches, but gets some type of summaries of pages, and it doesn't actually go into the links themselves to give me details.
Re: Audio is the one area small labs are winning
#94OpenAI and google are too scared of music industry lawyers to tackle this. Internally they without a doubt have models that would crush these startups over night if they chose to release them.
What about Disney's lawyers? GenAI for images exists ...
Re: Audio is the one area small labs are winning
#95I check every day for a new full-duplex model. I was so hyped about PersonaPlex from their demos, but in my test it was oddly dumb and unable to follow instructions. So I am hoping for something like PersonaPlex but a bit larger. Has anyone tested MiniCPM-o?.How is it at instruction following?
It's actively under development. Do you have a particular use-case in mind?
Re: Audio is the one area small labs are winning
#96Speaking of audio + AI, here's a "learning hack" I've been trying with voice mode, and the 3 big AI labs still haven't nailed it: While on a walk with mobile phone + earphones, dump an article/paper/HN-Post/github-repo into the mobile chat app (chat-gpt, claude or gemini), and use voice mode to have it walk you through it conversationally, so you can ask follow up questions during the walk-thru and the AI would do we…
Re: Audio is the one area small labs are winning
#97My understanding is that this is purely a strategic choice by the bigger labs. When OpenAI released Whisper, it was by far best-in-class, and they haven't released any major upgrades since then. It's been 3.5 years... Whisper is older than ChatGPT. Gemini 3 Pro Preview has superlative audio listening comprehension. If I send it a recording of myself in a car, with me talking, and another passenger talking to the driv…
Re: Audio is the one area small labs are winning
#98Earlier quoted context omitted.
What's the point of saying that without backing it up? Either you think it's so obvious it doesn't need backing up (in which case you don't need to say it), or ...? The reason it matters is that soon, any time somebody sees a comment they don't like or think is stupid, they'll just say, "eh a bot said that," and totally dilute the rest of the discussion, even if the comment was real.
We are *quickly* approaching a Tuesday where "bot detected, opinion rejected" is going to be a default assumption.
Re: Audio is the one area small labs are winning
#99Is there something that will read books to me? I.e I have some books in epub format and want audiobook versions for them, with a nice voice.
Re: Audio is the one area small labs are winning
#100Good article and I agree with everything in there. For my own voice agent I decided to make him PTT by default as the problems of the model accurately guessing the end of utterance are just too great. I think it can be solved in the future but, I haven't seen a really good example of it being done with modern day tech including this labs. Fundamentally it all comes down to the fact that different humans have differen…
Check out Sparrow-0. The demo shows an impressive ability to predict when the speaker has finished talking: https://www.tavus.io/post/sparrow-0-advancing-conversational...