Live data from Hacker News

Perceptually lossless (talking head) video compression at 22kbit/s

mlumiste.com

121–130 of 144 posts

Re: Perceptually lossless (talking head) video compression at 22kbit/s

#121
post #81

Earlier quoted context omitted.

Soft vs hard is based on how closely the world tracks with modern physics/science. As such even just FTL is soft, let alone everything else that doesn’t fit.

That is simply your personal definition, right? You don’t claim to be definitive?

It’s a classic definition. Soft/hard science fiction has two meanings either the topic is focused on hard sciences (physics) vs soft sciences (sociology) or “It can also refer to science fiction which prioritizes human emotions over scientific accuracy or plausibility.[1]”

So it’s not universal but is an accepted definition that any deviation from the possible or probable (for example, including faster-than-light travel or paranormal powers) to be a mark of "softness."

https://en.wikipedia.org/wiki/Soft_science_fiction

Popular science fiction is generally extremely soft, but occasionally you get stuff like The Cold Equations where the plot is driven by real world constraints. Even then it included FTL so a purest would call it soft.

Re: Perceptually lossless (talking head) video compression at 22kbit/s

#122

These sorts of models pop here quite a bit, and they ignore fundamental facts of video codecs (video specific lossy compression technologies). Traditional codecs have always focused on trade offs among encode complexity, decode complexity, and latency. Where complexity = compute. If every target device ran a 4090 at full power, we could go far below 22kbps with a traditional codec techniques for content like this. 22…

This is my field as well, although I come from the neural network angle. Learned video codecs definitely do look promising: Microsoft's DCVC-FM ( https://github.com/microsoft/DCVC ) beats H.267 in BD-rate. Another benefit of the learned approach is being able to run on soon commodity NPUs, without special hardware accommodation requirements. In the CLIC challenge, hybrid codecs (traditional + learned components) are…

Winning in bd rate though isn't hard. You need to win in bd rate and have a hardware implementable, power efficient, cheap decoder.

Agreed hybrid presents real opportunity.

Re: Perceptually lossless (talking head) video compression at 22kbit/s

#123
post #43

This reminds me of a scene in "A Fire Upon the Deep" (1992) where they're on a video call with someone on another spaceship; but something seems a bit "off". Then someone notices that the actual bitrate they're getting from the other vessel is tiny -- far lower than they should be getting given the conditions -- and so most of what they're seeing on their own screens isn't actual video feed, but their local computer'…

Was that the same book that had the concept of (paraphrasing using modern terminology) doing interstellar communications by sending back and forth LLMs trained on the people who wanted to talk, prompted to try and get a good business deal or whatever?

That happened in Redemption Ark by Alastair Reynolds (2002), though of course the idea may also have been used before or since.

Re: Perceptually lossless (talking head) video compression at 22kbit/s

#124
post #50
post #42

Earlier quoted context omitted.

Ability to tell MP3 from the original source was always dependent on encoder quality, bitrate, and the source material. In the mid 2000's, I tried to encode all of my music as MP3. Most of it sounded just fine because pop/rock/alt/etc are busy and "noisy" by design. But some songs (particularly with few instruments, high dynamic range, and female vocals) were just awful no matter how high I cranked the bitrate. And I…

These are always fun https://abx.digitalfeed.net/ https://www.npr.org/sections/therecord/2015/06/02/411473508/...

I find generic ABX tests not great, personally, because I generally don’t know what to be listening for. However, with songs I’ve listened to lossless my whole life, it’s much easier to spot encoding failures - an intuitive “wait, that cymbal crash sounded different” or “that multi-instrument harmonic should be cleaner/dirtier.”

That being said, 320Kbps AAC encoded by Core Audio I’ve found to be pretty much transparent with anything I’ve thrown at it. Anything less than that (256Kbps AAC, 320Kbps MP3, etc) I can ABX sometimes, as long as I’m familiar with the source material, and usually only with quality headphones. Although no streaming services provide that, so I’m stuck with ALAC through Apple Music for streaming (which is more convenient than my old solution, which was transcoding and transferring to an iPod ~20k songs selected from ~90k in my library based on a variety of rules than never gave me the song I’m looking for). And really, ~900Kbps lossless is pretty easy to justify these days with 5G data speeds and generally much higher data transfer limits.

The other downside to storing losing encodings these days is the fact that almost everyone uses Bluetooth for their listening, which is an additional lossy encoding. While 256Kbps AAC/320Kbps MP3 might be transparent in some cases, when it’s re-encoded it very rarely is (in my experience)

Re: Perceptually lossless (talking head) video compression at 22kbit/s

#125

These sorts of models pop here quite a bit, and they ignore fundamental facts of video codecs (video specific lossy compression technologies). Traditional codecs have always focused on trade offs among encode complexity, decode complexity, and latency. Where complexity = compute. If every target device ran a 4090 at full power, we could go far below 22kbps with a traditional codec techniques for content like this. 22…

Really? Would there be a way to replicate this with currently available encoders? I'd like to try it

Re: Perceptually lossless (talking head) video compression at 22kbit/s

#126

These sorts of models pop here quite a bit, and they ignore fundamental facts of video codecs (video specific lossy compression technologies). Traditional codecs have always focused on trade offs among encode complexity, decode complexity, and latency. Where complexity = compute. If every target device ran a 4090 at full power, we could go far below 22kbps with a traditional codec techniques for content like this. 22…

This is my field as well, although I come from the neural network angle. Learned video codecs definitely do look promising: Microsoft's DCVC-FM ( https://github.com/microsoft/DCVC ) beats H.267 in BD-rate. Another benefit of the learned approach is being able to run on soon commodity NPUs, without special hardware accommodation requirements. In the CLIC challenge, hybrid codecs (traditional + learned components) are…

Did you mean H.266? Or is there some secret H.267 that hasn't been agreed upon yet

Re: Perceptually lossless (talking head) video compression at 22kbit/s

#127
the only information that needs to be transmitted is the change in expression, pose and facial keypoints

Does anyone else remember the weirder (for lack of a better term) features of MPEG-4 part 2, like face and body animation? It did something like that, but as far as I know nearly no one used that feature for anything.

https://en.wikipedia.org/wiki/Face_Animation_Parameter

and in the worst, trust on the internet will be heavily undermined

...as long as the model doesn't include data to put a shoe on one's head.

Re: Perceptually lossless (talking head) video compression at 22kbit/s

#130
post #59

Earlier quoted context omitted.

Can you propose a better term for the concept then? Perceiving something as lossless is a real world metric that has a proper use case. "Perceptually lossless" does not try to imply that it is not lossy.

The term for this is "transparency." A codec is "transparent" if people can't tell the difference between the original and the compressed version.

I work in graphics. Calling this transparency would be a terrible idea and make a lot of discussions around compression of videos and images with actual transparency very confusing.

Does the compression algorithm work well for transparency? Yes, it's effect on transparency is totally transparent! In fact the transparency is fully transparently compressed by our codec.

Yeah, don't do this please. Perceptually lossless is a term I've heard lots of times before and companies developing codecs usually have a fairly strong technical basis for making the claim. As in, it's not like they just glance at the results and say "yep, looks good to me". Rather, they'll be looking and spectral curves and image diffs - probably also motion diffs for videos - and checking whether they the losses are small enough to be undetectable to human eyes.

Post reply on HN