The issue described here, to my understanding, is that you find or create csam and then manipulate it so that its fingerprint collides with another image that you want to be flagged as csam. You then submit the manipulated version of the found or generated image to the authority. First, at what point does the authority go "uh... Where did you get this from?" Practically speaking, the people doing this would have to b…
>Like if an ordinary citizen showed up with csam the excuse "oh I found this JPEG in a trunk in my late grandpa's attic" isn't gonna fly Just post it to 4chan, it’s actively monitored by intelligence services
Tainting the CSAM client-side scanning database
201–210 of 276 posts
Re: Tainting the CSAM client-side scanning database
#202This is basically what I suggested we do a few years ago if Apple added client-side content scanners. Seems all need to here is obtain a finger print of a CSAM image then you can reverse engineer a non CSAM image to match that finger print. Distribute this image wide enough and you effectively render these algorithms useless. Which is a shame because they could be used for good, but it seems kinda obvious this will e…
You can't do this without committing an illegal act (downloading CSAM) and it also wouldn't work (because there's a second server side hash), so no I don't think you should do this.
Re: Tainting the CSAM client-side scanning database
#203I am against the idea of scanning for the reason that the author pointed out: It's trivial to repurpose the technology to use it in dystopian ways. I however have precisely zero concerns about impersonating hashes: 1. It's trivial to deal with tainting the database: both secondary hashing and more invasive hashes deal with that problem. 2. It's trivial to deal with impersonated hashes, all positives can be scanned on…
If you can make one algorithm collide, you can make two collide.
The increase in difficulty is multiplicative not additive.
Re: Tainting the CSAM client-side scanning database
#204Earlier quoted context omitted.
I think that some people are terrified that their (possibly AI generated) "loli pictures" would be caught by a scanner.
What I'm concerned about is a system that flags me for a crime based on a database I can't audit based on mechanisms with an entirely too high false positive rate. Because the database can't be audited by anyone but a select group we have to trust that it only contains actual bad images. I do not trust that such databases don't also contain images that are embarrassing to powerful/connected people. I also do not trus…
e.g. there isn't a high false positive rate, the attack in this article doesn't work because the attacker doesn't have access to all the hashing algorithms used, and it doesn't text the police.
Re: Tainting the CSAM client-side scanning database
#205Earlier quoted context omitted.
This one happened just a month ago so it might be what the parent comment is referring to. Basically, somebody made a fake company to make bogus copyright claims against someone to hurt their channel. Youtube refuses to deliver counterclaims unless the YTer puts their government name on it (he originally tried to deliver it via an attorney). Additionally, the other party is actively trying to compromise the YTer's ot…
What I don’t get is this: why does youtube still have such a monopoly after all this time? It’s had a shitty reputation since I can remember, why don’t creators just band together and take their viewership somewhere less hostile?
Re: Tainting the CSAM client-side scanning database
#206Earlier quoted context omitted.
It's really about preventing images from circulating. Yes, the database maintainers would notice, but it could take them a while to get around to it. And they're not going to be eager to remove a collision, because that would effectively "legalize" the child porn member of the image pair. There are lots of images that could be useful to suppress temporarily . And if you actually succeed in suppressing the false-posit…
>Yes, the database maintainers would notice, but it could take them a while to get around to it. And they're not going to be eager to remove a collision, because that would effectively "legalize" the child porn member of the image pair. There are lots of images that could be useful to suppress temporarily. They may not know it immediately, but the actual CSAM image wouldn't actually be shared in any real numbers whic…
> Once again, this behavior would get Hungary booted pretty quickly.
I'm sorry; I should have made myself clearer. In that paragraph, I've taken an aside and moved from Hungary (or any other government) as the adversary to child-porn-sharers-in-general as the adversary, and also changed the adversary goal.
No matter what you do, you can't "boot" the child porn sharers, because they're the ones who actually define what images you're legitimately trying to block.
> Also it is important to remember these are hashes. Not all images of the US flag would trigger the system only that specific image that has a hash collision.
They're approximate perceptual hashes, designed to come up with close values on close pictures. The US flag has an officially defined appearance. You'll get the same hash for any two close-cropped, straight-on images of the flag, the kind you might embed in your Web site. They'll be at least as close as two reprocessed versions of the same child porn image.
I was oversimplifying, though. You're not going to be able to tweak just any child porn image to make it hash like just any flag, unless you're willing to distort it into unrecognizability. And flags and logos might be bad candidates in general, because they're going to give DCT output that's wildly different than what you'll get from most photos. But if you had a relatively large library of child porn and a relatively large library of heavily-used effectively unbannable images of whatever kind, you should be able to find a lot of the child porn images that you can tweak to hash like one or another of the heavily used ones.
... and if you're generating fake child porn from scratch using ML, you can probably hack your ML model to bake the hash of one or another unbannable image into everything it creates. You could probably make those matches pretty damned close.
So, once they got the whole thing down, they should be able to force the system to greatly tighten its match thresholds and/or deal with a really high false positive rate. In fact, they could probably make things bad enough that they could end up with a pretty large collection of child porn that the operators would be forced to completely exclude from detection.
Now back on the original authoritarian threat model:
> I just think this in a complicated and ineffectual bullet that can only be fired once because it has a fingerprint of the person who fired it.
For the "authoritarian" purpose, you may be right... although I wouldn't be surprised if it stayed under the radar for longer than you think. If the image you want to suppress only circulates among your own people, and if you're the authority who receives and verifies the reports on your own people, then all you have to do is to keep the overall volume down enough that it doesn't make anybody suspicious enough to demand that you show them the reports you're getting.
If you're a relatively small country, the number of people who share some local meme you care about may be quite a bit smaller than the number of people who share some new real child porn image.
Re: Tainting the CSAM client-side scanning database
#207If there were N algorithms, is it feasible to create B' image such that its fingerprint matches A in every algorithm?
This is a really stupid idea. Politicians who are demanding this are way out of their depths.
Re: Tainting the CSAM client-side scanning database
#208This is basically what I suggested we do a few years ago if Apple added client-side content scanners. Seems all need to here is obtain a finger print of a CSAM image then you can reverse engineer a non CSAM image to match that finger print. Distribute this image wide enough and you effectively render these algorithms useless. Which is a shame because they could be used for good, but it seems kinda obvious this will e…
> Seems all need to here is obtain a finger print of a CSAM image then you can reverse engineer a non CSAM image to match that finger print. Distribute this image wide enough and you effectively render these algorithms useless. You can't do this without committing an illegal act (downloading CSAM) and it also wouldn't work (because there's a second server side hash), so no I don't think you should do this.
Re: Tainting the CSAM client-side scanning database
#209It says: > This shows that the database can be tainted with non-CSAM material by an entity that can submit entries to it. Actually, it can easily be tainted by anybody . Take your massaged hash-colliding image, which remember is still visually child porn , and post it on some pedos-R-us forum. The people who maintain the database actively troll those forums. They'll see the image and add the hash to the database for…
[flagged]
Re: Tainting the CSAM client-side scanning database
#210It says: > This shows that the database can be tainted with non-CSAM material by an entity that can submit entries to it. Actually, it can easily be tainted by anybody . Take your massaged hash-colliding image, which remember is still visually child porn , and post it on some pedos-R-us forum. The people who maintain the database actively troll those forums. They'll see the image and add the hash to the database for…
You'd be publishing child porn, which I think is not the wisest thing to be doing.