Live data from Hacker News

Tainting the CSAM client-side scanning database

blog.xot.nl

111–120 of 276 posts

Re: Tainting the CSAM client-side scanning database

#111

Any computational method that relies on a function that converts m bits (an image) to n bits (a fingerprint) where m > n will always be vulnerable to such an attack. The smaller n is compared to m, the easier it is to counterfeit something with that signature. It is not new knowledge. The only way to be certain, unfortunately, is for a human to look at it. When the allegation is as serious as CSAM, I would rather be…

> Any computational method that relies on a function that converts m bits (an image) to n bits (a fingerprint) where m > n will always be vulnerable to such an attack. The "vulnerable" depends on your definition: SHA-2 and SHA-3 are both still quite safe against preimage attacks, and even second preimage attacks require significant work to pull off for SHA-2, and I am unaware of any meaningful second preimage attack…

These systems don't rely on cryptographic hashes but in fact the reverse: content sensitive hashes.

They essentially take an image and scale it to a small thumbnail. The values of all those reduced pixels are the hash of the original image. When a new image is scanned it's just doing a similarity check against the database of those hashes of known "bad" images. A hit triggers checking against an image hash performed with a separate algorithm. Hits against multiple hashes triggers a "bad image alarm" and ruins a person's life.

Changing a few input bits in an image doesn't usually change the perceptual hash because they're meant to be resistant to small amounts of localized noise. It is possible to add noise that will change a perceptual hash. It's also possible to manipulate an image so transforms (scaling etc) get wildly different results.

This leads to two exploits. The first is an attacker manipulates a "good" image such that when hashed it matches perceptual hashes of a "bad" image. The attacker then sends a bunch of these to a target triggering the "bad image alarm" and essentially SWATs the target. The second is to manipulate bad images in a recoverable way to get them pasted bad image scanners.

Re: Tainting the CSAM client-side scanning database

#112
post #108

Earlier quoted context omitted.

My ethical premise is that there is no direct victim of AI generated CSAM, but that it's worth criminalizing because otherwise it further victimizes victims of existing law. In other words, there is a societal victim of it. To me it's the same ethical premise but interpreted within two different frameworks: one that's purely idealistic, and one that's based in practical reality.

AI cannot generate CSAM, because AI cannot abuse children. AI makes fictional images, which definitionally cannot be images of child sexual abuse. There is literally no victim of any kind, even conceptually, in the case of computer generated imagery. It should be protected artistic expression.

What if a police officer generates some AI CSAM and then sells it to someone who thinks it's real? There's still "no victim," but the buyer thinks that there was. Are they guilty of a crime?

Your logic would seem to imply that there's no crime with possession of real CSAM either, and that the only crime lies with the original abuser who took the pictures.

Re: Tainting the CSAM client-side scanning database

#113
post #58

Earlier quoted context omitted.

How is providing a F-Droid repo the same as being restricted to only being available via F-Droid?

It was in reply to: > If push comes to shove they'll be fine and pressure to black box signal in the EU is unlikely to hold up if they can just move users to another app store .

Yes but the assumption is that "push comes to shove" and they have to choose between pulling it from the EU play store and installing black box MITM software.

Having relatively easy alternatives in place would reduce the leverage the EU has to actually practically enforce this and hopefully pressure them into at minimum non-enforcement and preferably walking back the obviously unenforceable legislation.

But if it looks like they could practically force out Signal and co or force them to adopt this MITM if they want to continue existing, then they might be more likely to pursue it.

Re: Tainting the CSAM client-side scanning database

#114

The issue described here, to my understanding, is that you find or create csam and then manipulate it so that its fingerprint collides with another image that you want to be flagged as csam. You then submit the manipulated version of the found or generated image to the authority. First, at what point does the authority go "uh... Where did you get this from?" Practically speaking, the people doing this would have to b…

The general public submitting CSAM directly would indeed be highly unlikely, but the scenario we need to consider involves those in positions of authority who can manipulate systems behind the scenes. Imagine that an unflattering or satirical image of Viktor Orban is circulating in France, and let's say it becomes viral, inciting discussions that the Hungarian government finds detrimental to its international image.…

Won't that easily be found out when the hash matches the image of Viktor Orban rather than an image of CSAM. I'm not sure the legal system is as stupid as you think it is. Then Hungary would just have their hashes reviewed.

Sure a conspiracy of all the relevant authorities across the whole EU would work, ... but that seems a stretch to enable what, political elites to _cooperatively_ censor images the public hold. There are easier ways, surely.

Re: Tainting the CSAM client-side scanning database

#115
This attack also works for people who are not in a position to submit entries to the database as you could very well plant the generated matching images on a server set up with a compromised credit card and report it yourself.

You could select pictures from targets publically available photos or more insidiously look to compromised accounts cloud storage and generate fake offensive images that match peoples actual baby pictures whether for harm or blackmail.

Re: Tainting the CSAM client-side scanning database

#118
post #2

Because client-side scanning is not going to work, and no one wants to government issued black box binary to send their conversations and photos to unnamed police person randomly, the non-compliance is the only way. People just start to use chat programs in the EU that do not comply. This would be Signal, Telegram, others. The EU can fine and fight with Meta/WhatsApp, Apple, others, but that’s about it. The EU bureau…

[deleted]

Re: Tainting the CSAM client-side scanning database

#119

Earlier quoted context omitted.

So you have to plant the manipulated csam somewhere that it'll be found by law enforcement and added to the database and just hope you did a good enough job to not be tracked?

Look if I want to host some white noise I can't be held responsible for what happens if people xor it together with some other file hosted elsewhere.

Erm, of course you can, you are then just distributing means for acquiring CSAM, or taking part in conspiracy to distribute it.

Re: Tainting the CSAM client-side scanning database

#120

Earlier quoted context omitted.

You'd be publishing child porn, which I think is not the wisest thing to be doing.

A spy agency or secret police doing this to use the client-side scanning system as a surveillance/censorship system for political media would not care about that.

Even if such systems were only used for their stated purpose of finding CSAM, because no one can audit the database of hashes there's no guarantee that non-CSAM images aren't in there.

Just being accused of having CSAM is a life ruining event. Even if someone is eventually cleared of charges their life is forever altered. The "We Got Him!!" headlines are front page news, retractions are filed in a basement filing cabinet with a sign that says "Beware of Leopard".

Post reply on HN