Live data from Hacker News

Tainting the CSAM client-side scanning database

blog.xot.nl

121–130 of 276 posts

Re: Tainting the CSAM client-side scanning database

#121

Earlier quoted context omitted.

I assume the answer to that will be that there is no need to differentiate between them. And honestly, I agree with that argument. Possession of CSAM should be illegal regardless of whether it's "real" or not. But the proposed scanning system is the wrong solution, regardless of any "real or AI" ambiguity, because it's possible to generate false positives with nonsense images that aren't even close to the expected CS…

> I assume the answer to that will be that there is no need to differentiate between them. And honestly, I agree with that I disagree. The point is to reduce actual child abuse. The images are in a way only tangential. If an image is made with an AI with no actual child being abused, then it shouldn't be a crime. In a way, it's better , because it will distract the crowd of people into this sort of stuff from activit…

>> I assume the answer to that will be that there is no need to differentiate between them. And honestly, I agree with that

>I disagree. The point is to reduce actual child abuse.

There are limited resources, practically the only way to do this is to make it illegal to have anything that looks real (or looks derived from a real situation, in a 'I will know it when I see it way'). Otherwise, you're just making an almost impassable defence of 'it is fake' or 'I thought it was fake'. Then you can't practically reduce actual child abuse.

Re: Tainting the CSAM client-side scanning database

#122

Earlier quoted context omitted.

[flagged]

Angle brackets are traditionally used for actual quotes. > Step one, break cryptographically secure hashing system These "perceptual hashes" are not (and cannot be) cryptographically secure, and practical collision attacks have been tested and published . Did you actually read the main article at all? I don't think I'm even going to bother with the rest.

PhotoDNA doesn't use CS hashes, but others I've looked at do use them. But you're right, you probably don't wanna respond to my parody of you suggestion to trade in CSAM...

Re: Tainting the CSAM client-side scanning database

#123

Earlier quoted context omitted.

You'd be publishing child porn, which I think is not the wisest thing to be doing.

Tainting the database is probably already chargeable as something . The point is that the person doing this doesn't expect to get caught. And frankly they're probably right in that.

> Tainting the database is probably already chargeable as something

This is definitely true. Good old “intentionally accesses a computer without authorization or exceeds authorized access” in the CFAA. It’s so vague as to be able to make almost _anything_ the government doesn’t like involving computers illegal.

Re: Tainting the CSAM client-side scanning database

#124

Earlier quoted context omitted.

A spy agency or secret police doing this to use the client-side scanning system as a surveillance/censorship system for political media would not care about that.

Even if such systems were only used for their stated purpose of finding CSAM, because no one can audit the database of hashes there's no guarantee that non-CSAM images aren't in there. Just being accused of having CSAM is a life ruining event. Even if someone is eventually cleared of charges their life is forever altered. The "We Got Him!!" headlines are front page news, retractions are filed in a basement filing cab…

Which is why names and pictures of the accused should remain secret until the verdict, like it already happens in many other countries. But that's a separate topic.

Re: Tainting the CSAM client-side scanning database

#125
post #80

This is basically what I suggested we do a few years ago if Apple added client-side content scanners. Seems all need to here is obtain a finger print of a CSAM image then you can reverse engineer a non CSAM image to match that finger print. Distribute this image wide enough and you effectively render these algorithms useless. Which is a shame because they could be used for good, but it seems kinda obvious this will e…

1. You mixed up generating child porn images that collide with targeted non-child-porn images, in order to taint the database and thereby suppress the target images, with generating non-child-porn images that collide with child porn images, in order to overwhelm the system with false positives. 2. You brought in matters that have nothing to do with hashing or child porn, but tend to generate flame wars. 3. You put sn…

Thank you.

1. Oh okay. I wasn't very clear there. By, "this is basically what I suggested we do a few years ago" I meant that we should look to exploit the limitations of these fingering printing algorithms to render them impractical. But yeah, my suggestion was that we should do the reverse. Rather than taint the database, taint the content.

2. That wasn't my intention. I was suggesting that there are likely other reasons governments are interested in this technology than, "think of the children". Hence, why I think we should seek to ways to render these algorithms ineffective were they deployed. But yes, please feel free to downvote if you feel my comment is unproductive. On reflection I agree it might be. I'm not really adding much here

3. You know what's annoying? I did that because I didn't to make an assertion either way to be as neutral as possible. I'm using quotes because others deem it racist, and that it's irrelevant what my personal position on that is. If you want my position on whether racism exists though, obviously I believe it does – I'd urge you to just assume good faith in the future.

Re: Tainting the CSAM client-side scanning database

#126

Earlier quoted context omitted.

Look if I want to host some white noise I can't be held responsible for what happens if people xor it together with some other file hosted elsewhere.

Erm, of course you can, you are then just distributing means for acquiring CSAM, or taking part in conspiracy to distribute it.

Under this logic any image at all on the entire internet is now "conspiracy to distribute" CSAM, since you can make an image diff from anything

Re: Tainting the CSAM client-side scanning database

#127
post #80

This is basically what I suggested we do a few years ago if Apple added client-side content scanners. Seems all need to here is obtain a finger print of a CSAM image then you can reverse engineer a non CSAM image to match that finger print. Distribute this image wide enough and you effectively render these algorithms useless. Which is a shame because they could be used for good, but it seems kinda obvious this will e…

I consider anyone promoting client side scanning of my media promoting a form of non-consensual violence against me.

Huh? I'm not promoting it?

I do think there may be a time and a place for them though... On school and work devices, for example.

That's just my weakly held opinion though. Isn't an opinion I would seek to "promote", just one I wouldn't object to it.

Re: Tainting the CSAM client-side scanning database

#128

Earlier quoted context omitted.

Seems like you want to bring all the success of the war on drugs to the war on generated images.

No, I don't support any automated scanning system or really any sort of "going out of our way" to find new criminals. The reason I think it's a bad idea to differentiate between real or AI is similar to the arguments against "means testing" for distributing benefits. You don't want to put real victims in a situation where they're deprived of justice because they can't prove that their victimization was "real." Imagin…

>> Imagine a real CSAM criminal claiming a defense that they "thought it was AI generated."

Saying "I thought this heroin was fake" is not a defense when caught with a bag of heroin, I don't see how this would be any different. It's not a magic out for anyone.

Re: Tainting the CSAM client-side scanning database

#129
post #45

Earlier quoted context omitted.

> Any computational method that relies on a function that converts m bits (an image) to n bits (a fingerprint) where m > n will always be vulnerable to such an attack. The "vulnerable" depends on your definition: SHA-2 and SHA-3 are both still quite safe against preimage attacks, and even second preimage attacks require significant work to pull off for SHA-2, and I am unaware of any meaningful second preimage attack…

But then bypassing the filter would just require to change a few bits here and there.

No the whole purpose of these “hashes” is that they’re robust to that. The attack model they’re designed for is image manipulation to avoid the hash match, not manipulating manipulating images to trigger the hash.

There are numerous papers on doing bit manipulation to cause miss classification.

Re: Tainting the CSAM client-side scanning database

#130

The issue described here, to my understanding, is that you find or create csam and then manipulate it so that its fingerprint collides with another image that you want to be flagged as csam. You then submit the manipulated version of the found or generated image to the authority. First, at what point does the authority go "uh... Where did you get this from?" Practically speaking, the people doing this would have to b…

The general public submitting CSAM directly would indeed be highly unlikely, but the scenario we need to consider involves those in positions of authority who can manipulate systems behind the scenes. Imagine that an unflattering or satirical image of Viktor Orban is circulating in France, and let's say it becomes viral, inciting discussions that the Hungarian government finds detrimental to its international image.…

>The Hungarian government - who presumably has access to the EU CSAM database (or can coerce those who do), might attempt to add a fingerprint of a manipulated CSAM image that collides with the fingerprint of the satirical image.

Then what? What does that achieve? There would be a huge spike in images identified as CSAM which would obviously throw up red flags. It seems like this would mostly just be headache for the law enforcement across the EU. It isn't like France is going to arrest a huge number of people without any investigation or thought of how this one CSAM image spread so far so quickly. And if we are talking about the result in Hungary, that government doesn't need this tool to abuse its power. Why go through all that effort? They could just do the equivalent of rubber-hose cryptanalysis.

Post reply on HN