>That's why most codecs have packet loss concealment (PLC) that can fill in for missing packets with plausible audio that just extrapolates what was being said and avoids leaving a hole in the audio ...How far can ML PLC "hallucinate" audio? A sound , a syllable, a whole word, half a sentence? Can I trust anymore what I hear?
Opus 1.5 released: Opus gets a machine learning upgrade
121–130 of 151 posts
Re: Opus 1.5 released: Opus gets a machine learning upgrade
#122Some people hyping it as AGI on social media
Re: Opus 1.5 released: Opus gets a machine learning upgrade
#123They’ll have my upvote just for writing ML instead AI. Seriously, this is very exciting developments for audio compression.
Machine Learning is Artificial Intelligence. Just look at Wikipedia: https://en.wikipedia.org/wiki/Artificial_intelligence
Re: Opus 1.5 released: Opus gets a machine learning upgrade
#124The main limitation for such codecs is CPU/battery life - and I like how they sparsely applied ML in it here and there, combining it with classic approach (non-ML algos) to achieve better tradeoff of CPU vs quality. E.g. for better low bitrate support/LACE - "we went for a different approach: start with the tried-and-true postfilter idea and sprinkle just enough DNN magic on top of it." The key was not to feed raw au…
It's a really smart application of ML: helping around the edges and not letting the ML algo invent pheonems or even whole words by accident. ML transcription has a similar trade-off of performing better on some benchmarks but also hallucinating results.
I still think the pendulum will swing back there again, to have even better battery/larger models on mobile.
Re: Opus 1.5 released: Opus gets a machine learning upgrade
#125I'm using Opus as one of the main codecs in my peer-to-peer audio streaming library ( https://git.iem.at/cm/aoo/ - still alpha), so this is very exciting news! I'll definitely play around with these new ML features!
Interesting :)
Re: Opus 1.5 released: Opus gets a machine learning upgrade
#126Earlier quoted context omitted.
Why is the ethics question important? It is a new feature for an audio codec, not a new material to teach in your kids curriculum.
This is a great question! Here's a related failure case that I think illustrates the issue. In my country, public restroom facilities replaced all the buttons and levers on faucets, towel dispensers, etc. with sensors that detect your hand under the faucet. Black people tell me they aren't able to easily use these restrooms. I was surprised when I heard this, but if you google this, it's apparently a thing. Why does…
Re: Opus 1.5 released: Opus gets a machine learning upgrade
#127Earlier quoted context omitted.
At least in the U.S., anyone can look up all the patents a person/entity owns. So, your fraud wouldn’t get very far. https://assignment.uspto.gov/patent/index.html#/patent/searc...
"I represent the holders of the patents in question" is simple enough. I wonder if it's fraud, if all you're putting out is an unverifiable claim on the net. The pool operators do that all the time.
It’s like someone selling a deed to land they don’t own, or leasing a property they don’t own to a tenant.
Re: Opus 1.5 released: Opus gets a machine learning upgrade
#128Earlier quoted context omitted.
The BT SIG moves kind of slow and there's a really long tail of devices. Until there's a chip with native Opus support (that's as cheap as ones with AAC etc) you wouldn't get Opus support even if it was in the spec. Realistically for most headphones people actually buy AAC (LC and HE) is more than good enough encoding quality for the audio the headphones can produce. Even if Opus was in the spec and Opus-supporting c…
They chose to make totally new inferior LC3 codec though. Also, on my system (Android phone + BTR5/BTR15 Bluetooth DAC + Sennheiser H600) all options sound realy crappy compared to plain old usb, everything else is the same. LDAC 990kbps is less crappy, by sheer brute force. I suspect it's not only codec but other co-factors as well (like mandatory DSP on phone side)
This is hearsay: Bluetooth SIG considered Opus but rejected it because it was computationally too expensive. This came out of the hearing aid group, where battery life and complexity are a major restriction.
So when you compare codecs in this space, the metric you want to look at is quality vs. CPU cycles. In that regard LC3 outperforms many contemporary codecs.
Regarding sound quality it's simply a matter of setting the appropriate bitrate. So if Opus is transparent at 150 kbps, and LC3 at 250 kbps thats totally acceptable if that gives you more battery life.
Re: Opus 1.5 released: Opus gets a machine learning upgrade
#129Earlier quoted context omitted.
"I represent the holders of the patents in question" is simple enough. I wonder if it's fraud, if all you're putting out is an unverifiable claim on the net. The pool operators do that all the time.
Someone who has no basis to bring a patent infringement claim selling a settlement of such a claim to an alleged infringer is clearly fraud. It’s like someone selling a deed to land they don’t own, or leasing a property they don’t own to a tenant.
Few people will even bother to try, and if so, you keep them at a distance with some random bullshit (communication can break down _sooo_ easily), but it certainly poisons the well by adding a layer of FUD to the tech you're targeting with your claims.
Standard fare of patent pool operators, and it's high time to reciprocate.
Re: Opus 1.5 released: Opus gets a machine learning upgrade
#130Earlier quoted context omitted.
Here's a question that should have the same/similar answer: Increasingly some part of the job interviews is being handled over the internet. All other things being equal, people are likely to have a more positive response to candidates with more pleasant voice. So if new ML-enhanced codecs become more common, we may find that some group X has a just slightly worse quality score than others. Over enough samples that w…
I don't think it's a given that we shouldn't keep using that codec. For example, maybe the improvement is due to an open source hacker working in their spare time to make the world a better place. Do we tell them their contribution isn't welcome until it meets the community's benchmark for equity? Your same argument can also be used to degrade the performance for all other groups, so that group X isn't unfairly disad…
If improving the codec is a matter of training on dataset A vs dataset B, that’s an easier change.