Live data from Hacker News

Researchers tested AI watermarks and broke all of them

wired.com

41–50 of 91 posts

Re: Researchers tested AI watermarks and broke all of them

#41
post #34

Earlier quoted context omitted.

the blu-ray is encrypted and the player requires a key to decrypt it if you want your blu-ray player to be able to read blu-ray disks you sign a contract that says you will respect the revocation list if you change your mind later: your player key will be revoked and new blu-rays won't play on your device (it's actually more sophisticated than this... they can block specific players too)

That doesn't quite sound right to me. I don't think a new Blu-ray disc could be released that continues to be readable by some old readers but is no longer readable by other old readers.

> I don't think a new Blu-ray disc could be released that continues to be readable by some old readers but is no longer readable by other old readers.

you can obviously think whatever you want, but you'd be completely wrong

DVD supported this 20 years ago, blu-ray's system is far more sophisticated and can even block individual players

    The approach of AACS provisions each individual player with a unique set of decryption keys which are used in a broadcast encryption scheme. This approach allows licensors to "revoke" individual players, or more specifically, the decryption keys associated with the player. Thus, if a given player's keys are compromised and published, the AACS LA can simply revoke those keys in future content, making the keys/player useless for decrypting new titles.
(from https://en.wikipedia.org/wiki/Advanced_Access_Content_System)

the spec also supports a persistent CRL so a new disk can also stop your old disks from working

Re: Researchers tested AI watermarks and broke all of them

#42

Earlier quoted context omitted.

>Of course, Intel's HDCP master key was either leaked or reverse-engineered, so anyone can generate their own valid device keys Of an older version of HDCP. New media can require a higher HDCP version where that bypass isn't possible.

That's certainly possible, but did that actually happen in practice? The downside is obviously that you essentially start selling new Blu-ray discs that don't work on any old Blu-ray players. I feel like I would have heard if that happened. Unless maybe the old players could issue firmware updates?

>That's certainly possible, but did that actually happen in practice?

Yes, see 4k blurays for requiring HDCP 2.2 requiring people to get a new bluray player.

Re: Researchers tested AI watermarks and broke all of them

#44
post #6

It’s like captcha, highly annoying to users and authors, but if you don’t want to pay it works against low spend bots

I'm pretty certain the article is saying it is _not_ like captcha in that it is so trivial to circumvent that it's completely useless, rather than just useless sometimes.

Re: Researchers tested AI watermarks and broke all of them

#45

Earlier quoted context omitted.

Interesting. I don't understand the revocation process though. What stops the blu-ray reader from just ignoring the revocation list on the disk?

The revocation list is for the TV. Intel revokes a TV's key, distributes the updated revocation list on new Blu-ray discs, and when a compliant Blu-ray player is playing one of those new discs it will refuse to negotiate with a revoked TV. Now that I think of it, I wonder if compliant Blu-ray players actually save the new revocation entries and then continue refusing to negotiate with revoked TVs even for old Blu-ray…

If so, couldn't a malicious disc revoke all TVs?

Re: Researchers tested AI watermarks and broke all of them

#46
post #8

For written text, the problem may be even harder. Identifying the human author of text is a field called "stylometry" but this result shows that some simple transformations reduce the success to random chance [1]. Similarly, I suspect watermarking LLM output is probably unworkable. The output of a smart model could be de-watermarked by fine tuning a dumb open source model on the initial output, and then regenerating…

The idea of telling a human generated "the quick brown fox..." from a machine-generated one was always a fantasy. Text has no birthmark.

Current LLMs have stylistic quirks imprinted on them by RLHF (ChatGPT's endless "it should be noted" and "it is important to remember that" verbiage is a good example), but they learned those from human writing.

Re: Researchers tested AI watermarks and broke all of them

#47
post #17
post #10

We need to focus on the other direction. How can we have chains of trust for content creation, such as for real video. Content can be faked, but not necessarily easily faked from the same sources that make use of cryptographic signing. The attacks can sign the own work, so you'd need ways to distinguish those cases, but device level keys, organizational keys, distribution keys all can provide provenance chains that c…

I was thinking the other day about embedding keys in cameras, etc. but came up with the problem that you could just wire up a computer that BEHAVES like a CCD sensor and send whatever the hell you feel like in to the signing hardware, so you feed in your fake image and it gets signed by the camera as though it were real. I assume smarter people than me have put much more time into the problem, so I'd be interested to…

a very large number of consumer and high-end camera equipment does have unique ID of the device. Some metadata stores those device IDs along with other values like mfg or light settings. Its not signed into a binding chain, only marked at image creation time. The vast majority of people I expect would just copy the data files blindly.

Re: Researchers tested AI watermarks and broke all of them

#49

Wasn't this obvious from the get go that this can't work? If AI will eventually generate say 10k by 10k images, I can resize to 2.001k by 1.999k or similar, and I just don't get how any subtle signal in the pixels can persist through that. Maybe you could do something at the compositional level, but that seems restrictive to the output. Maybe something about like larger regions average color balance or something? But…

> Wasn't this obvious from the get go that this can't work?

It needs to publicly fail first to manufacture consent for full surveillance of every human interaction with any computer. Nobody would ever want that otherwise.

Re: Researchers tested AI watermarks and broke all of them

#50
post #10

We need to focus on the other direction. How can we have chains of trust for content creation, such as for real video. Content can be faked, but not necessarily easily faked from the same sources that make use of cryptographic signing. The attacks can sign the own work, so you'd need ways to distinguish those cases, but device level keys, organizational keys, distribution keys all can provide provenance chains that c…

What's the threat vector you're trying to mitigate here? If you're wondering whether a movie that claims to be produced by Disney really was, if it's in theaters or on Disney+, then you can trust it was actually made by Disney or at least licensed to them. As long as the Washington Post still employs its own photographers and doesn't accept imagery submission from the general public, you should be able to trust a photo published in the Washington Post is at least of something real, unless you just don't trust the Post itself. If you're thinking a YouTube channel or something, unless the channel got hacked, seemingly anything published there was really published by the channel owner. Maybe they're showing you something made by AI that isn't real, but as the owner of their own signing key, nothing would prevent them from signing an AI-generated image.

If you're talking someone on Twitter or Facebook is putting a photo in your feed claiming a human photographed BLM throwing bricks through a window, don't trust shit being posted on Facebook or Twitter no matter what, probably, but even there, unless the profile was hacked, you either trust the person who owns it or you don't. Nothing would prevent them from signing a forgery of reality that they legitimately forged themselves. Even with device-level keys, what are you trying to prove? You can pay actors to throw bricks through windows.

I guess the concern is this doesn't scale as well as asking Midjourney to do it, but I wonder to what extent that is even true. With 8 billion people on the planet and counting and a whole lot of them doing this shit, given the limited input bandwidth of human sense organs, there is seemingly some maximum saturation of bullshit a person can be exposed to that a lot of people have already hit, and having the Internet host even more of it doesn't mean they'll grow bigger eyes and a faster brain that can actually ingest more bullshit than it already does.

Post reply on HN