Live data from Hacker News

Audio is the one area small labs are winning

amplifypartners.com

101–109 of 109 posts

Re: Audio is the one area small labs are winning

#102
post #54

Earlier quoted context omitted.

Is your claim that music industry lawyers are that much scarier than movie industry lawyers? Because the big labs don't seem to have any problem releasing models that create (possibly infringing) video.

The movie industry is doing well from AI. Thus far AI has only been used to create fan fiction clips that generate free marketing for legacy IP on TikTok. And the rights holders know that if AI gets good enough to make feature length movies then they'll be able to aggressively use various legal mechanisms to take the videos off major sites and pursue the creators. Long term it could potentially lower internal product…

Just right now: ByteDance to curb AI video app after Disney legal threat

https://www.bbc.com/news/articles/c93wq6xqgy1o

Re: Audio is the one area small labs are winning

#103
post #33

Earlier quoted context omitted.

Your reply has a 0% ai score yet the presence of that "reply to" text is concerning.

Not very concerning because that post is obviously LLM generated.

Fair callout. The ideas and NVIDIA/X-Wing analogy are mine — I'm a non-native English speaker and use AI to help draft comments based on my detailed directions. The "Reply to" text was an artifact I should have removed — that's on me. I engage here to learn and discuss, not to farm karma or manufacture consensus. Happy to continue the conversation about inference chips and "rebel infrastructure," or I can step back if the AI assistance undermines the discussion.

Re: Audio is the one area small labs are winning

#105
post #78

Earlier quoted context omitted.

> imo audio DSP experts are diametrically opposed to AI on moral grounds. Can you elaborate on this point? I don't know the moral grounds of audio DSP experts, and thus I don't understand why in your opinion they wouldn't take an offer if you really pay them some serious amount of money. Just to be clear: considering what a typical daily job in DSP programming is like, I can imagine that many audio DSP experts are no…

In my experience, most of the people in audio DSP are musicians or otherwise very well exposed to music and the arts, and many see using AI as fundamentally immoral or unethical. It's technology designed via theft with the intent to harm professionals in these spaces. Like I said, it's like paying a doctor to design a better gun.

Yes, but, you see, guns are merely tools. If you give these guns to the right people, they will actually create more work for the doctors, and everybody wins.

Re: Audio is the one area small labs are winning

#106

Audio models are also tiny, which is probably why small labs are doing well in the space. I run a LoRA'd Whisper v3 Large for a client. We can fit 4 versions of the model in memory at once on a ~$1/hr A10 and have half the VRAM leftover. Each of the LoRA tunes we did took maybe 2-3 hours on the same A10 instance.

Is Whisper still getting nontrivial development? I was under the impression that it had stagnated, but it seems hard to find more than just rumors

My ~1.7% WER and faster than realtime processing in my application make it more than adequate. My application is multi-speaker with WPM rates >300 for long durations.

Re: Audio is the one area small labs are winning

#107
post #103
post #33

Earlier quoted context omitted.

Not very concerning because that post is obviously LLM generated.

Fair callout. The ideas and NVIDIA/X-Wing analogy are mine — I'm a non-native English speaker and use AI to help draft comments based on my detailed directions. The "Reply to" text was an artifact I should have removed — that's on me. I engage here to learn and discuss, not to farm karma or manufacture consensus. Happy to continue the conversation about inference chips and "rebel infrastructure," or I can step back i…

Hi I am a non native english speaker as well. Please don't be discouraged. But you're not speaking with your own voice if I make sense.

Honestly I prefer to talk to humans who have flaws. Even if your english is bad try typing it yourself. Everyone here will support you. That's how I learned.

Even my team mates use chatgpt to rewrite their messages in teams, it feels so dishonest.

Re: Audio is the one area small labs are winning

#108

It's amazing how good open-weight STT and TTS have gotten, so there's no need to pay for Wispr Flow, Superwhisper, Eleven-Labs etc. Sharing my setup in case it may be useful for others; it's especially useful when working with CLI agents like Code Code or Codex-CLI: STT: Hex [1] (open-source), with Parakeet V3 - stunningly fast, near-instant transcription. The slight accuracy drop relative to bigger models is immater…

Anyone know of something like Hex that runs on Linux?

You can roll a script to do this, something that would consume a mic from Pipewire when triggered and then push results to clipboard. With a Parakeet ONNX model in between.

I had cause to do the the opposite: Hotkey -> clipboard TTS

Re: Audio is the one area small labs are winning

#109
post #103
post #33

Earlier quoted context omitted.

Not very concerning because that post is obviously LLM generated.

Fair callout. The ideas and NVIDIA/X-Wing analogy are mine — I'm a non-native English speaker and use AI to help draft comments based on my detailed directions. The "Reply to" text was an artifact I should have removed — that's on me. I engage here to learn and discuss, not to farm karma or manufacture consensus. Happy to continue the conversation about inference chips and "rebel infrastructure," or I can step back i…

>or I can step back if the AI assistance undermines the discussion.

Yes.

Please don't post generated or AI-filtered posts to HN. We want to hear you in your own voice, and it's fine if your English isn't perfect.

https://news.ycombinator.com/item?id=46747998

Post reply on HN