Live data from Hacker News

Audio is the one area small labs are winning

amplifypartners.com

71–80 of 109 posts

Re: Audio is the one area small labs are winning

#71

Audio models are also tiny, which is probably why small labs are doing well in the space. I run a LoRA'd Whisper v3 Large for a client. We can fit 4 versions of the model in memory at once on a ~$1/hr A10 and have half the VRAM leftover. Each of the LoRA tunes we did took maybe 2-3 hours on the same A10 instance.

Is Whisper still getting nontrivial development? I was under the impression that it had stagnated, but it seems hard to find more than just rumors

Re: Audio is the one area small labs are winning

#72

It's amazing how good open-weight STT and TTS have gotten, so there's no need to pay for Wispr Flow, Superwhisper, Eleven-Labs etc. Sharing my setup in case it may be useful for others; it's especially useful when working with CLI agents like Code Code or Codex-CLI: STT: Hex [1] (open-source), with Parakeet V3 - stunningly fast, near-instant transcription. The slight accuracy drop relative to bigger models is immater…

Anyone know of something like Hex that runs on Linux?

Re: Audio is the one area small labs are winning

#73

for a laugh enter nonsense at https://gradium.ai/ You get all kinds of weird noises and random words. Jack is often apologetic about the problem you are having with the Hyperion xt5000 smart hub.

Hey Rob. I'm not on the tech team here at Gradium (I do GTM) but still curious where you found the glitch? Were you entering words into the STT in the bottom of the front page? Can you share an example so I can replicate? Many thanks!

Re: Audio is the one area small labs are winning

#74

It's amazing how good open-weight STT and TTS have gotten, so there's no need to pay for Wispr Flow, Superwhisper, Eleven-Labs etc. Sharing my setup in case it may be useful for others; it's especially useful when working with CLI agents like Code Code or Codex-CLI: STT: Hex [1] (open-source), with Parakeet V3 - stunningly fast, near-instant transcription. The slight accuracy drop relative to bigger models is immater…

Anyone know of something like Hex that runs on Linux?

Handy is cross-platform, including linux

Re: Audio is the one area small labs are winning

#75

Earlier quoted context omitted.

Anyone know of something like Hex that runs on Linux?

Handy is cross-platform, including linux

+1 for Handy, it's very easy to get running and once it is you don't have to think about it again.

Re: Audio is the one area small labs are winning

#76
post #30

Can someone reccomend to me: a service that will generate a loopable engine drone for a "WWII Plane Japan Kawasaki Ki-61"? It doesn't have to be perfect, just convincing in a hollywood blockbuster context, and not just a warmed over clone of a Merlin engine sound. Turns out Suno will make whatever background music I need, but I want a "unique sound effect on demand" service. I'm not convinced voice AI stuff is sustai…

A fun alternative could be using a physically based engine sound synthesizer - for example - https://github.com/Engine-Simulator/engine-sim-community-edi...

Re: Audio is the one area small labs are winning

#77
post #7

Earlier quoted context omitted.

True, but there's a fun irony: the Rebels' X-Wings are powered by GPUs from a company that's... checks relationships ...also supplying the Empire. NVIDIA's basically the galaxy's most successful arms dealer, selling to both sides while convincing everyone they're just "enabling innovation." The real rebels would be training audio models on potato-patched RP2040s. Brave souls, if they exist.

The company behind the T-65B X-Wing, Incom Corporation, did supply the Empire, as they did the Republic Navy before. By 0 BBY, Incom was nationalized by the Imperials. The X-Wing became the mainstay Alliance fighter because the plans were stolen by some defecting Incom engineers.

wonder why the Empire never ran any black x wings

Re: Audio is the one area small labs are winning

#78
post #21

Earlier quoted context omitted.

imo audio DSP experts are diametrically opposed to AI on moral grounds. Good luck hiring the good ones. It's like paying doctors to design guns.

> imo audio DSP experts are diametrically opposed to AI on moral grounds. Can you elaborate on this point? I don't know the moral grounds of audio DSP experts, and thus I don't understand why in your opinion they wouldn't take an offer if you really pay them some serious amount of money. Just to be clear: considering what a typical daily job in DSP programming is like, I can imagine that many audio DSP experts are no…

In my experience, most of the people in audio DSP are musicians or otherwise very well exposed to music and the arts, and many see using AI as fundamentally immoral or unethical. It's technology designed via theft with the intent to harm professionals in these spaces.

Like I said, it's like paying a doctor to design a better gun.

Re: Audio is the one area small labs are winning

#79
I check every day for a new full-duplex model. I was so hyped about PersonaPlex from their demos, but in my test it was oddly dumb and unable to follow instructions.

So I am hoping for something like PersonaPlex but a bit larger.

Has anyone tested MiniCPM-o?.How is it at instruction following?

Re: Audio is the one area small labs are winning

#80
post #79

I check every day for a new full-duplex model. I was so hyped about PersonaPlex from their demos, but in my test it was oddly dumb and unable to follow instructions. So I am hoping for something like PersonaPlex but a bit larger. Has anyone tested MiniCPM-o?.How is it at instruction following?

It's actively under development. Do you have a particular use-case in mind?
Post reply on HN