Live data from Hacker News

Lyra audio codec enables high-quality voice calls at 3 kbps bitrate

cnx-software.com

1–10 of 207 posts

Re: Lyra audio codec enables high-quality voice calls at 3 kbps bitrate

#6
post #2

Maybe in the future, all we need is a speech example, some AI and the continious transmission of text for low data voice transmission?

I think this will be the rough direction, but not exactly text, rather some other efficient, machine-readable embedding of speech that is also able to carry tone and rhythm effectively and pronounciation accurately and unambiguously.

Re: Lyra audio codec enables high-quality voice calls at 3 kbps bitrate

#7
post #6
post #2

Maybe in the future, all we need is a speech example, some AI and the continious transmission of text for low data voice transmission?

I think this will be the rough direction, but not exactly text, rather some other efficient, machine-readable embedding of speech that is also able to carry tone and rhythm effectively and pronounciation accurately and unambiguously.

Why speak then?

Re: Lyra audio codec enables high-quality voice calls at 3 kbps bitrate

#8
Nothing about licensing or patents. I assume the worst (read: unusable for small businesses)?

10+ years ago I worked in a small voip shop, where we had very high quality (low jitter), but low bandwidth connection. I researched many codecs of the time (2010-ish).

We liked speex, because it can be used "without strings attached". Also, I can choose the quality depending on the bandwidth. Although for low bandwidth g729 was better. Which we couldn't use because of royalties (but allowed myself to test it).

We chose alaw/ulaw when bandwidth was not a concern, and speex when it was.

Since it does not mention usability outside of google, I also find this comparison unfair or incomplete: if you are comparing a proprietary codec, compare it to g729. If you are comparing a codec to speex, it should be open/free.

Edit: grammar

Re: Lyra audio codec enables high-quality voice calls at 3 kbps bitrate

#10
When Google's announcement [1] was posted a few days ago, I listened to their samples and heard an odd effect in the "chocolate bread" sample (the video chat example) [1], which is not mirrored in this article.

On that sample, I felt [2] that the Lyra version exaggerates the pronunciation of the phrase 'with chocolate' in a way that meaningfully differs from the speaker's original. It weakens the voiced 'th' to nothingness, and overshoots both the lead consonant and first vowel of 'choc', and then proceeds to wash the entire rest of the sentence with a peculiar brightened voice that's high, lacks consonant definition, and is close to ringing.

I'm guessing it's actually style transfer, because though the result sounds not much like the speaker's original, the result is reminiscent of the speech pattern and accent that people with East Asian and Southeast Asian ancestry adopt when speaking American English. It was surprising, given that the speaker doesn't sound like that in the original. I wonder if others hear this too.

While Lyra sounds richer and wider-band than Opus or Speex at these bitrates, the degradations and artifacts of those codecs are universally recognized (through years of familiarity with telephones) as compression artifacts and not innate features of the speaker themselves. Therefore listeners can be expected to be sympathetic to the quality issues and not attribute the whole of the sound on the speaker's person.

If AI-trained voice synthesizer codecs become the norm, and it performs well on most speakers, that expectation will go away, and the resulting audio will be attributed wholly to the speaker. That increases the impact of mistakes and misrepresentations introduced by the codec, unbeknowst to the speaker and listener.

[1] https://ai.googleblog.com/2021/02/lyra-new-very-low-bitrate-...

[2] https://news.ycombinator.com/item?id=26282519

Post reply on HN