Live data from Hacker News

Audio is the one area small labs are winning

amplifypartners.com

91–100 of 109 posts

Re: Audio is the one area small labs are winning

#91
post #89

It's amazing how good open-weight STT and TTS have gotten, so there's no need to pay for Wispr Flow, Superwhisper, Eleven-Labs etc. Sharing my setup in case it may be useful for others; it's especially useful when working with CLI agents like Code Code or Codex-CLI: STT: Hex [1] (open-source), with Parakeet V3 - stunningly fast, near-instant transcription. The slight accuracy drop relative to bigger models is immater…

Is Hex MacOS only?

Yes

Re: Audio is the one area small labs are winning

#93
Speaking of audio + AI, here's a "learning hack" I've been trying with voice mode, and the 3 big AI labs still haven't nailed it:

While on a walk with mobile phone + earphones, dump an article/paper/HN-Post/github-repo into the mobile chat app (chat-gpt, claude or gemini), and use voice mode to have it walk you through it conversationally, so you can ask follow up questions during the walk-thru and the AI would do web-search etc. I know I could do something like this with NotebookLM, but I want to engage in the conversation, and NotebookLM does have interactive mode but it has been super-flaky to say the least.

I pay for ChatGPT Pro and the voice mode is really bad: it pretends to do web searches and makes up things, and when pushed says it didn't actually read the article. Also the voice sounds super-condescending.

Gemini Pro mobile app - similarly refuses to open links and sounds as if it's talking to a baby.

Claude mobile app was the best among these - the voice is very tolerable in terms of tone, but like the others it can't open links. I does do web searches, but gets some type of summaries of pages, and it doesn't actually go into the links themselves to give me details.

Re: Audio is the one area small labs are winning

#94
post #66
post #47

OpenAI and google are too scared of music industry lawyers to tackle this. Internally they without a doubt have models that would crush these startups over night if they chose to release them.

What about Disney's lawyers? GenAI for images exists ...

Disney is actually quite excited about GenAI [0]

[0] https://openai.com/index/disney-sora-agreement/

Re: Audio is the one area small labs are winning

#95
post #79

I check every day for a new full-duplex model. I was so hyped about PersonaPlex from their demos, but in my test it was oddly dumb and unable to follow instructions. So I am hoping for something like PersonaPlex but a bit larger. Has anyone tested MiniCPM-o?.How is it at instruction following?

It's actively under development. Do you have a particular use-case in mind?

Outgoing phone calls.

Re: Audio is the one area small labs are winning

#96

Speaking of audio + AI, here's a "learning hack" I've been trying with voice mode, and the 3 big AI labs still haven't nailed it: While on a walk with mobile phone + earphones, dump an article/paper/HN-Post/github-repo into the mobile chat app (chat-gpt, claude or gemini), and use voice mode to have it walk you through it conversationally, so you can ask follow up questions during the walk-thru and the AI would do we…

I have found that the "advanced voice mode" is dumb as a box of rocks compared to their "basic" TTS version, so I disable it. I've switched to Claude, so I don't know if that's still an option, but if you are tied to ChatGPT, definitely disable it if possible!

Re: Audio is the one area small labs are winning

#97
post #40

My understanding is that this is purely a strategic choice by the bigger labs. When OpenAI released Whisper, it was by far best-in-class, and they haven't released any major upgrades since then. It's been 3.5 years... Whisper is older than ChatGPT. Gemini 3 Pro Preview has superlative audio listening comprehension. If I send it a recording of myself in a car, with me talking, and another passenger talking to the driv…

Nvidia released Parakeet which claimed superiority. Doesn't negate your point but I did want to add it.

Re: Audio is the one area small labs are winning

#98
post #43

Earlier quoted context omitted.

What's the point of saying that without backing it up? Either you think it's so obvious it doesn't need backing up (in which case you don't need to say it), or ...? The reason it matters is that soon, any time somebody sees a comment they don't like or think is stupid, they'll just say, "eh a bot said that," and totally dilute the rest of the discussion, even if the comment was real.

We are *quickly* approaching a Tuesday where "bot detected, opinion rejected" is going to be a default assumption.

Yeah I thought I knew what "post-truth" meant back in 2016 but boy was I wrong.

Re: Audio is the one area small labs are winning

#100
post #65

Good article and I agree with everything in there. For my own voice agent I decided to make him PTT by default as the problems of the model accurately guessing the end of utterance are just too great. I think it can be solved in the future but, I haven't seen a really good example of it being done with modern day tech including this labs. Fundamentally it all comes down to the fact that different humans have differen…

Check out Sparrow-0. The demo shows an impressive ability to predict when the speaker has finished talking: https://www.tavus.io/post/sparrow-0-advancing-conversational...

Thanks, ill read it now.
Post reply on HN