Live data from Hacker News

Audio is the one area small labs are winning

amplifypartners.com

51–60 of 109 posts

Re: Audio is the one area small labs are winning

#51
post #47

OpenAI and google are too scared of music industry lawyers to tackle this. Internally they without a doubt have models that would crush these startups over night if they chose to release them.

I'm not sure it's just fear of lawyers, although that's definitely part of it. Big companies have way more to lose reputationally and legally, so the bar for releasing something is much higher

Re: Audio is the one area small labs are winning

#52

Good article and I agree with everything in there. For my own voice agent I decided to make him PTT by default as the problems of the model accurately guessing the end of utterance are just too great. I think it can be solved in the future but, I haven't seen a really good example of it being done with modern day tech including this labs. Fundamentally it all comes down to the fact that different humans have differen…

It feels like this is one of those areas where the last 10% of polish will take 90% of the effort

Re: Audio is the one area small labs are winning

#53
Moshi was an amazing tech demo, building the entire stack from scratch in 6 months with a small team was an amazing show of skill: 7B text LLM data + training, emotive TTS for synth data generation (again model + data collection), synth data pipeline, novel speech codec, rust inference stack for low latency, audio LLM architecture incl. text "thoughts" stream which was novel.

But, this piece is a fluff piece: "underfunded" means a total of around $400 million ($330 million in the initial round, $70 million for Gradium). Compare to Elevenlabs who used a $2 million pre-seed for creating their initial product.

A bunch of other stuff there is disingenuous, like comparing their 7B model to Llama-3 405B (hint: the 7B model is a _lot_ dumber). There's also the outright lie: team of 4 made Moshi, which is corrected _in the same piece_ to 8 if you read enough.

Re: Audio is the one area small labs are winning

#54
post #47

OpenAI and google are too scared of music industry lawyers to tackle this. Internally they without a doubt have models that would crush these startups over night if they chose to release them.

Is your claim that music industry lawyers are that much scarier than movie industry lawyers? Because the big labs don't seem to have any problem releasing models that create (possibly infringing) video.

Re: Audio is the one area small labs are winning

#55
post #54
post #47

OpenAI and google are too scared of music industry lawyers to tackle this. Internally they without a doubt have models that would crush these startups over night if they chose to release them.

Is your claim that music industry lawyers are that much scarier than movie industry lawyers? Because the big labs don't seem to have any problem releasing models that create (possibly infringing) video.

> Is your claim that music industry lawyers are that much scarier than movie industry lawyers?

Not qoez:

You have to balance market opportunities with the risk of reputational damage and litigation risk.

Video will probably make a lot more money than audio, so you are willing to take a bigger risk. Additionally, at least for Google there exists a strong synergy between their video generation models and YouTube, which makes it even more sensible for Google to make video models available to the public despite these risks.

Re: Audio is the one area small labs are winning

#56
post #21
post #6

[flagged]

imo audio DSP experts are diametrically opposed to AI on moral grounds. Good luck hiring the good ones. It's like paying doctors to design guns.

> imo audio DSP experts are diametrically opposed to AI on moral grounds.

Can you elaborate on this point? I don't know the moral grounds of audio DSP experts, and thus I don't understand why in your opinion they wouldn't take an offer if you really pay them some serious amount of money.

Just to be clear: considering what a typical daily job in DSP programming is like, I can imagine that many audio DSP experts are not the best culture fit for AI companies, but this doesn't have anything to do with morality.

Re: Audio is the one area small labs are winning

#57
post #54
post #47

OpenAI and google are too scared of music industry lawyers to tackle this. Internally they without a doubt have models that would crush these startups over night if they chose to release them.

Is your claim that music industry lawyers are that much scarier than movie industry lawyers? Because the big labs don't seem to have any problem releasing models that create (possibly infringing) video.

well i guess the music industry is a lot more monopolized than video, plus there is a lot of video out there that isn't "movies," while there's not a lot of music that isn't... "music"

Re: Audio is the one area small labs are winning

#58

OpenAI being the death star and audio AI being the rebels is such a weird comparison, like what? Wouldn't the real rebels be the ones running their own models locally?

Audio AI companies are just another death star, intent on reducing human creativity to "make a song like Let it be, but in the style of Eminem, and change the lyrics to match the birthday of my mother in law". The only rebels are musicians resisting this hedge-fund driven monstrosity.

Re: Audio is the one area small labs are winning

#59

Moshi was an amazing tech demo, building the entire stack from scratch in 6 months with a small team was an amazing show of skill: 7B text LLM data + training, emotive TTS for synth data generation (again model + data collection), synth data pipeline, novel speech codec, rust inference stack for low latency, audio LLM architecture incl. text "thoughts" stream which was novel. But, this piece is a fluff piece: "underf…

Stopped reading there: "This model (Moshi) could [...] recite an original poem in a French accent (research shows poems sound better this way)."

Re: Audio is the one area small labs are winning

#60
post #54
post #47

OpenAI and google are too scared of music industry lawyers to tackle this. Internally they without a doubt have models that would crush these startups over night if they chose to release them.

Is your claim that music industry lawyers are that much scarier than movie industry lawyers? Because the big labs don't seem to have any problem releasing models that create (possibly infringing) video.

The movie industry is doing well from AI.

Thus far AI has only been used to create fan fiction clips that generate free marketing for legacy IP on TikTok. And the rights holders know that if AI gets good enough to make feature length movies then they'll be able to aggressively use various legal mechanisms to take the videos off major sites and pursue the creators. Long term it could potentially lower internal production costs by getting rid of actors & writers.

Music is very different. The production cost is already zero, and people generating their own Taylor Swift songs is a real competitive threat to Spotify etc.

Post reply on HN