Is there something that will read books to me? I.e I have some books in epub format and want audiobook versions for them, with a nice voice.
Audio is the one area small labs are winning
101–109 of 109 posts
Re: Audio is the one area small labs are winning
#102Earlier quoted context omitted.
Is your claim that music industry lawyers are that much scarier than movie industry lawyers? Because the big labs don't seem to have any problem releasing models that create (possibly infringing) video.
The movie industry is doing well from AI. Thus far AI has only been used to create fan fiction clips that generate free marketing for legacy IP on TikTok. And the rights holders know that if AI gets good enough to make feature length movies then they'll be able to aggressively use various legal mechanisms to take the videos off major sites and pursue the creators. Long term it could potentially lower internal product…
Re: Audio is the one area small labs are winning
#103Earlier quoted context omitted.
Your reply has a 0% ai score yet the presence of that "reply to" text is concerning.
Not very concerning because that post is obviously LLM generated.
Re: Audio is the one area small labs are winning
#104Re: Audio is the one area small labs are winning
#105Earlier quoted context omitted.
> imo audio DSP experts are diametrically opposed to AI on moral grounds. Can you elaborate on this point? I don't know the moral grounds of audio DSP experts, and thus I don't understand why in your opinion they wouldn't take an offer if you really pay them some serious amount of money. Just to be clear: considering what a typical daily job in DSP programming is like, I can imagine that many audio DSP experts are no…
In my experience, most of the people in audio DSP are musicians or otherwise very well exposed to music and the arts, and many see using AI as fundamentally immoral or unethical. It's technology designed via theft with the intent to harm professionals in these spaces. Like I said, it's like paying a doctor to design a better gun.
Re: Audio is the one area small labs are winning
#106Audio models are also tiny, which is probably why small labs are doing well in the space. I run a LoRA'd Whisper v3 Large for a client. We can fit 4 versions of the model in memory at once on a ~$1/hr A10 and have half the VRAM leftover. Each of the LoRA tunes we did took maybe 2-3 hours on the same A10 instance.
Is Whisper still getting nontrivial development? I was under the impression that it had stagnated, but it seems hard to find more than just rumors
Re: Audio is the one area small labs are winning
#107Earlier quoted context omitted.
Not very concerning because that post is obviously LLM generated.
Fair callout. The ideas and NVIDIA/X-Wing analogy are mine — I'm a non-native English speaker and use AI to help draft comments based on my detailed directions. The "Reply to" text was an artifact I should have removed — that's on me. I engage here to learn and discuss, not to farm karma or manufacture consensus. Happy to continue the conversation about inference chips and "rebel infrastructure," or I can step back i…
Honestly I prefer to talk to humans who have flaws. Even if your english is bad try typing it yourself. Everyone here will support you. That's how I learned.
Even my team mates use chatgpt to rewrite their messages in teams, it feels so dishonest.
Re: Audio is the one area small labs are winning
#108It's amazing how good open-weight STT and TTS have gotten, so there's no need to pay for Wispr Flow, Superwhisper, Eleven-Labs etc. Sharing my setup in case it may be useful for others; it's especially useful when working with CLI agents like Code Code or Codex-CLI: STT: Hex [1] (open-source), with Parakeet V3 - stunningly fast, near-instant transcription. The slight accuracy drop relative to bigger models is immater…
Anyone know of something like Hex that runs on Linux?
I had cause to do the the opposite: Hotkey -> clipboard TTS
Re: Audio is the one area small labs are winning
#109Earlier quoted context omitted.
Not very concerning because that post is obviously LLM generated.
Fair callout. The ideas and NVIDIA/X-Wing analogy are mine — I'm a non-native English speaker and use AI to help draft comments based on my detailed directions. The "Reply to" text was an artifact I should have removed — that's on me. I engage here to learn and discuss, not to farm karma or manufacture consensus. Happy to continue the conversation about inference chips and "rebel infrastructure," or I can step back i…
Yes.
Please don't post generated or AI-filtered posts to HN. We want to hear you in your own voice, and it's fine if your English isn't perfect.