Live data from Hacker News

Opus 1.5 released: Opus gets a machine learning upgrade

opus-codec.org

121–130 of 151 posts

Re: Opus 1.5 released: Opus gets a machine learning upgrade

#121
post #82

>That's why most codecs have packet loss concealment (PLC) that can fill in for missing packets with plausible audio that just extrapolates what was being said and avoids leaving a hole in the audio ...How far can ML PLC "hallucinate" audio? A sound , a syllable, a whole word, half a sentence? Can I trust anymore what I hear?

You never can when lossy compression is involved. It is commonly considered good practice to verify that the communication partner understood what was said, e.g., by restating, summarizing, asking for clarification, follow-up questions etc.

Re: Opus 1.5 released: Opus gets a machine learning upgrade

#123
post #4

They’ll have my upvote just for writing ML instead AI. Seriously, this is very exciting developments for audio compression.

Machine Learning is Artificial Intelligence. Just look at Wikipedia: https://en.wikipedia.org/wiki/Artificial_intelligence

So are compilers and interpreters. The terminology changes, but since we still don't have a general, systematic, and precise definition of what "intelligence" means, the term AI is and always was ill-founded and a buzzword for investors. Sometimes, people get disillusioned, and that's how you get AI winters.

Re: Opus 1.5 released: Opus gets a machine learning upgrade

#124
post #43

The main limitation for such codecs is CPU/battery life - and I like how they sparsely applied ML in it here and there, combining it with classic approach (non-ML algos) to achieve better tradeoff of CPU vs quality. E.g. for better low bitrate support/LACE - "we went for a different approach: start with the tried-and-true postfilter idea and sprinkle just enough DNN magic on top of it." The key was not to feed raw au…

It's a really smart application of ML: helping around the edges and not letting the ML algo invent pheonems or even whole words by accident. ML transcription has a similar trade-off of performing better on some benchmarks but also hallucinating results.

fwiw, in ASR/speech transcription world, it looks reverse to me - in the past, there was lots of custom non-ML code & separate ML models for audio modeling and language modeling - but current SOTA ASRs are all e2e, and that's what's used even in mobile applications, iiuc.

I still think the pendulum will swing back there again, to have even better battery/larger models on mobile.

Re: Opus 1.5 released: Opus gets a machine learning upgrade

#125

I'm using Opus as one of the main codecs in my peer-to-peer audio streaming library ( https://git.iem.at/cm/aoo/ - still alpha), so this is very exciting news! I'll definitely play around with these new ML features!

> peer-to-peer audio streaming library

Interesting :)

Re: Opus 1.5 released: Opus gets a machine learning upgrade

#126
post #66

Earlier quoted context omitted.

Why is the ethics question important? It is a new feature for an audio codec, not a new material to teach in your kids curriculum.

This is a great question! Here's a related failure case that I think illustrates the issue. In my country, public restroom facilities replaced all the buttons and levers on faucets, towel dispensers, etc. with sensors that detect your hand under the faucet. Black people tell me they aren't able to easily use these restrooms. I was surprised when I heard this, but if you google this, it's apparently a thing. Why does…

I get your point, but in the example used - and I can think of couple others that start with "X replaced all controls with touch/voice/ML" - the much bigger ethical question is why did they do it in the first place. The new solution may or may not be biased differently than the old one, but it's usually inferior to the previous ones or simpler alternatives.

Re: Opus 1.5 released: Opus gets a machine learning upgrade

#127
post #88

Earlier quoted context omitted.

At least in the U.S., anyone can look up all the patents a person/entity owns. So, your fraud wouldn’t get very far. https://assignment.uspto.gov/patent/index.html#/patent/searc...

"I represent the holders of the patents in question" is simple enough. I wonder if it's fraud, if all you're putting out is an unverifiable claim on the net. The pool operators do that all the time.

Someone who has no basis to bring a patent infringement claim selling a settlement of such a claim to an alleged infringer is clearly fraud.

It’s like someone selling a deed to land they don’t own, or leasing a property they don’t own to a tenant.

Re: Opus 1.5 released: Opus gets a machine learning upgrade

#128

Earlier quoted context omitted.

The BT SIG moves kind of slow and there's a really long tail of devices. Until there's a chip with native Opus support (that's as cheap as ones with AAC etc) you wouldn't get Opus support even if it was in the spec. Realistically for most headphones people actually buy AAC (LC and HE) is more than good enough encoding quality for the audio the headphones can produce. Even if Opus was in the spec and Opus-supporting c…

They chose to make totally new inferior LC3 codec though. Also, on my system (Android phone + BTR5/BTR15 Bluetooth DAC + Sennheiser H600) all options sound realy crappy compared to plain old usb, everything else is the same. LDAC 990kbps is less crappy, by sheer brute force. I suspect it's not only codec but other co-factors as well (like mandatory DSP on phone side)

"Inferior" is relative. The main focus of LC3 was, as the name suggests, complexity.

This is hearsay: Bluetooth SIG considered Opus but rejected it because it was computationally too expensive. This came out of the hearing aid group, where battery life and complexity are a major restriction.

So when you compare codecs in this space, the metric you want to look at is quality vs. CPU cycles. In that regard LC3 outperforms many contemporary codecs.

Regarding sound quality it's simply a matter of setting the appropriate bitrate. So if Opus is transparent at 150 kbps, and LC3 at 250 kbps thats totally acceptable if that gives you more battery life.

Re: Opus 1.5 released: Opus gets a machine learning upgrade

#129
post #88

Earlier quoted context omitted.

"I represent the holders of the patents in question" is simple enough. I wonder if it's fraud, if all you're putting out is an unverifiable claim on the net. The pool operators do that all the time.

Someone who has no basis to bring a patent infringement claim selling a settlement of such a claim to an alleged infringer is clearly fraud. It’s like someone selling a deed to land they don’t own, or leasing a property they don’t own to a tenant.

The point isn't to sell a settlement. It's to publish "oh, we're totally serious that there are patents in our control. We won't tell you which ones, we don't tell you what we want from you. If you're interested, reach out to us."

Few people will even bother to try, and if so, you keep them at a distance with some random bullshit (communication can break down _sooo_ easily), but it certainly poisons the well by adding a layer of FUD to the tech you're targeting with your claims.

Standard fare of patent pool operators, and it's high time to reciprocate.

Re: Opus 1.5 released: Opus gets a machine learning upgrade

#130
post #100

Earlier quoted context omitted.

Here's a question that should have the same/similar answer: Increasingly some part of the job interviews is being handled over the internet. All other things being equal, people are likely to have a more positive response to candidates with more pleasant voice. So if new ML-enhanced codecs become more common, we may find that some group X has a just slightly worse quality score than others. Over enough samples that w…

I don't think it's a given that we shouldn't keep using that codec. For example, maybe the improvement is due to an open source hacker working in their spare time to make the world a better place. Do we tell them their contribution isn't welcome until it meets the community's benchmark for equity? Your same argument can also be used to degrade the performance for all other groups, so that group X isn't unfairly disad…

One thing the small mom-and-pop hacker types can do is disclose where bias can enter the system or evaluate it on standard benchmarks so folks can get an idea where it works and where it fails. That was the intent behind the top-level comment asking about bias, I think.

If improving the codec is a matter of training on dataset A vs dataset B, that’s an easier change.

Post reply on HN