Live data from Hacker News

Perceptually lossless (talking head) video compression at 22kbit/s

mlumiste.com

51–60 of 144 posts

Re: Perceptually lossless (talking head) video compression at 22kbit/s

#51
post #28

Earlier quoted context omitted.

GP is correct, that’s the definition of “lossy”. We don’t need to invent ever new marketing buzzwords for well-established technical concepts.

GP is incorrect. There is "Is identical", "looks identical" and "has lost sufficient detail to clearly not be the original." - being able to differentiate between these three states is useful.

Lossless means "is identical".

The other two are variations of lossy.

Calling one of them "perceptually lossless" is cheating, to the disadvantage of algorithms that honestly advertise themselves as lossy while still achieving "looks identical" compression.

Re: Perceptually lossless (talking head) video compression at 22kbit/s

#52

> But one overlooked use case of the technology is (talking head) video compression. > On a spectrum of model architectures, it achieves higher compression efficiency at the cost of model complexity. Indeed, the full LivePortrait model has 130m parameters compared to DCVC’s 20 million. While that’s tiny compared to LLMs, it currently requires an Nvidia RTX 4090 to run it in real time (in addition to parameters, a lar…

The trade-off may not be worth it today, but the processing power we can expect in the coming years will make this accessible to ordinary consumers. When your laptop or phone or AR headset has the processing power to run these models, it will make more efficient use of limited bandwidth, even if more bandwidth is available. I don't think available bandwidth will scale at the same rate as processing power, but even if it does, the picture be that much more realistic.

Re: Perceptually lossless (talking head) video compression at 22kbit/s

#53

This is very impressive, but “perceptually lossless” isn’t a thing and doesn’t make sense. It means “lossy”.

why not? if you change one pixel by one pixel brightness unit it is perceptually the same. for the record, I found liveportrait to be well within the uncanny valley. it looks great for ai generated avatars, but the difference is very perceptually noticeable on familiar faces. still it's great.

For one, it doesn't obey the transitive property like a truly lossless process should: unless it settles into a fixed point, a perceptually lossless copy of a copy of a copy, etc., will eventually become perceptually different. E.g., screenshot-of-screenshot chains, each of which visually resembles the previous one, but which altogether make the original content unreadable.

Re: Perceptually lossless (talking head) video compression at 22kbit/s

#54
post #28

Earlier quoted context omitted.

GP is correct, that’s the definition of “lossy”. We don’t need to invent ever new marketing buzzwords for well-established technical concepts.

GP is incorrect. There is "Is identical", "looks identical" and "has lost sufficient detail to clearly not be the original." - being able to differentiate between these three states is useful.

Importantly the first one is parameterless, but the second and third are parameterized by the audience. For example humans don't see colour very well, some animals have much better colour gamut, while some can't distinguish colour at all.

Re: Perceptually lossless (talking head) video compression at 22kbit/s

#55
post #28

Earlier quoted context omitted.

why not? if you change one pixel by one pixel brightness unit it is perceptually the same. for the record, I found liveportrait to be well within the uncanny valley. it looks great for ai generated avatars, but the difference is very perceptually noticeable on familiar faces. still it's great.

GP is correct, that’s the definition of “lossy”. We don’t need to invent ever new marketing buzzwords for well-established technical concepts.

But this marketing term has been regularly used in academic papers for nearly 50 years (or probably more), so it seems like it should get a pass IMO.

It's also used in the first paragraph of the Wikipedia article on the term "transparency" as it relates to data compression.

Re: Perceptually lossless (talking head) video compression at 22kbit/s

#56

Earlier quoted context omitted.

"AI" videos make many errors all the time, but most people are not aware of what to look for... Undetectable CGI is done in film/games all the time, and indeed it takes talent to hide the fact it is fake. One could rely on the media encoder to garble output enough to look more plausible (people on potato devices are used to looking at garbage content.) However, at the end of the day the "uncanny valley" effect takes…

Undetectable CGI in games ... right. I don't think you are a gamer. In movies it can be done with enough of manual tweaking by artists and a lot of photographic content around to borrow sense of reality from it. "Potato" devices by which I assume you mean average phones, currently have better resolutions than PCs had very recently and a lot still do (1080p). And a photo on 480p still looks more real than anything CGI…

I think most "AI" slop content falls under this phenomena:

https://www.youtube.com/watch?v=vJG698U2Mvo

Several 8bit games had their own aesthetic charm, but were at least fun...

Cheers, =3

Re: Perceptually lossless (talking head) video compression at 22kbit/s

#57
post #51

Earlier quoted context omitted.

GP is incorrect. There is "Is identical", "looks identical" and "has lost sufficient detail to clearly not be the original." - being able to differentiate between these three states is useful.

Lossless means "is identical". The other two are variations of lossy. Calling one of them "perceptually lossless" is cheating, to the disadvantage of algorithms that honestly advertise themselves as lossy while still achieving "looks identical" compression.

It's a well established term, though. It's been used in academic works for a long time (since at least 1970), and it's basically another term for the notion of "transparency" as it relates to data compression.

Re: Perceptually lossless (talking head) video compression at 22kbit/s

#58
I like how the saddle in the background moves with the reconstructed head; it probably works better with uncluttered backgrounds.

This is interesting tech, and the considerations in the introduction are particularly noteworthy. I never considered the possibility of animating 2D avatars with no 3D pipeline at all.

Re: Perceptually lossless (talking head) video compression at 22kbit/s

#59

This is very impressive, but “perceptually lossless” isn’t a thing and doesn’t make sense. It means “lossy”.

Can you propose a better term for the concept then? Perceiving something as lossless is a real world metric that has a proper use case. "Perceptually lossless" does not try to imply that it is not lossy.

Re: Perceptually lossless (talking head) video compression at 22kbit/s

#60
post #35

This is very impressive, but “perceptually lossless” isn’t a thing and doesn’t make sense. It means “lossy”.

Yeah, all lossy compression could be called "perceptually lossless" if the perception is bad enough...

A family member of mine didn't see the point of 1080p. Turned out they needed cataract surgery and got fancy replacement lenses in their eyes. After that, they saw the point.
Post reply on HN