Live data from Hacker News

Perceptually lossless (talking head) video compression at 22kbit/s

mlumiste.com

11–20 of 144 posts

Re: Perceptually lossless (talking head) video compression at 22kbit/s

#12

Earlier quoted context omitted.

Probably just disappointed at the wasted bandwidth: 24fps * 52 facial 3D marker * 16bit packed delta planar projected offsets (x,y) = 19.968 kbps And this is done in Unreal games on a potato graphics card all the time: https://apps.apple.com/us/app/live-link-face/id1495370836 I am sure calling modern heuristics "AI" gets people excited, but it doesn't seem "Magical" when trivial implementations are functionally equiv…

I think the point here is to make it photorealistic which everything apart from AI still fails at superhard.

Take a minute to look something up first, and then formulate a more interesting opinion for us to discuss:

https://www.unrealengine.com/en-US/metahuman

The artifacts in raster image data is nowhere near what a reasonable model can achieve even at low resolutions. =3

Re: Perceptually lossless (talking head) video compression at 22kbit/s

#14
> But one overlooked use case of the technology is (talking head) video compression.

> On a spectrum of model architectures, it achieves higher compression efficiency at the cost of model complexity. Indeed, the full LivePortrait model has 130m parameters compared to DCVC’s 20 million. While that’s tiny compared to LLMs, it currently requires an Nvidia RTX 4090 to run it in real time (in addition to parameters, a large culprit is using expensive warping operations). That means deploying to edge runtimes such as Apple Neural Engine is still quite a ways ahead.

It’s very cool that this is possible, but the compression use case is indeed .. a bit far fetched. A insanely large model requiring the most expensive consumer GPU to run on both ends and at the same time being limited in bandwidth so much (22kbps) is a _very_ limited scenario.

Re: Perceptually lossless (talking head) video compression at 22kbit/s

#15

Earlier quoted context omitted.

I think the point here is to make it photorealistic which everything apart from AI still fails at superhard.

Take a minute to look something up first, and then formulate a more interesting opinion for us to discuss: https://www.unrealengine.com/en-US/metahuman The artifacts in raster image data is nowhere near what a reasonable model can achieve even at low resolutions. =3

I know metahuman. As impressive as it is, when you judge by the standards of game graphics, if you are ever mislead into thinking metahumans are real humans or even real physically existing things it's time to see your eye doctor (and/or do MRI head scan).

On the other hand AI videos can be easily mistaken for people or hyper realistic physical sculptures.

https://img-9gag-fun.9cache.com/photo/aYQ776w_460svvp9.webm

There's something basic about how light works that traditional computer graphics still fails to grasp. Looking at its productions and comparing it to what AI generates is like looking at output of amateur and an artist. Sure, maybe artist doesn't always draw all 5 fingers but somehow captures the essence of the image in seemingly random arrangement of light and dark strokes, while amateur just tries to do their best but fails in some very significant ways.

Re: Perceptually lossless (talking head) video compression at 22kbit/s

#16

This is very impressive, but “perceptually lossless” isn’t a thing and doesn’t make sense. It means “lossy”.

why not? if you change one pixel by one pixel brightness unit it is perceptually the same.

for the record, I found liveportrait to be well within the uncanny valley. it looks great for ai generated avatars, but the difference is very perceptually noticeable on familiar faces. still it's great.

Re: Perceptually lossless (talking head) video compression at 22kbit/s

#19

> But one overlooked use case of the technology is (talking head) video compression. > On a spectrum of model architectures, it achieves higher compression efficiency at the cost of model complexity. Indeed, the full LivePortrait model has 130m parameters compared to DCVC’s 20 million. While that’s tiny compared to LLMs, it currently requires an Nvidia RTX 4090 to run it in real time (in addition to parameters, a lar…

130m parameters isn’t insanely large, even for smartphone memory. The high GPU usage is a barrier at the moment, but I wouldn’t put it past Apple to have 4090-level GPU performance in an iPhone before 2030.

Re: Perceptually lossless (talking head) video compression at 22kbit/s

#20

Earlier quoted context omitted.

Take a minute to look something up first, and then formulate a more interesting opinion for us to discuss: https://www.unrealengine.com/en-US/metahuman The artifacts in raster image data is nowhere near what a reasonable model can achieve even at low resolutions. =3

I know metahuman. As impressive as it is, when you judge by the standards of game graphics, if you are ever mislead into thinking metahumans are real humans or even real physically existing things it's time to see your eye doctor (and/or do MRI head scan). On the other hand AI videos can be easily mistaken for people or hyper realistic physical sculptures. https://img-9gag-fun.9cache.com/photo/aYQ776w_460svvp9.webm T…

"AI" videos make many errors all the time, but most people are not aware of what to look for... Undetectable CGI is done in film/games all the time, and indeed it takes talent to hide the fact it is fake.

One could rely on the media encoder to garble output enough to look more plausible (people on potato devices are used to looking at garbage content.) However, at the end of the day the "uncanny valley" effect takes over every-time even for live action data in a auto-generated asset, as the missing data can't be "Magically" recovered with 100% certainty.

Bye =3

Post reply on HN