Live data from Hacker News

Tainting the CSAM client-side scanning database

blog.xot.nl

211–220 of 276 posts

Re: Tainting the CSAM client-side scanning database

#211

The article considers "an entity that is allowed to propose new entries to the CSAM database". You don't even need this! You could target a whole "social cluster" of people without having any special privileges within this system. As an example, lets say you want to attack environmental protesters. For image A, you create a meme about climate change. For image B, you procure something that looks, to humans, like CSAM…

> and if the platform fulfills its obligations, that report will bubble up to the relevant authorities, who will review it and add its fingerprint to the database. I don't think NCMEC adds random images they find to the A1 list without knowing their origin. > until the average meme-savvy environmental protester's device gets flagged for further scrutiny. Who is this an attack on? A moderation contractor maybe, but no…

NCMEC’s database is smaller and infrequently updated, but Facebook’s database is much larger and more actively updated. This is also one of the reasons that Facebook/Meta produces many more reports than other providers who use the NCMEC database. Governments view the limited reach of the NCMEC database as a problem, and are encouraging providers to use more complete databases and to develop tools to detect novel CSAM (this is included in the EU proposed regulations.)

“Further scrutiny” here means your images are reported to a provider, which (at that point) means your provider has a list of reports that defeat the guarantees of E2E encryption. The threat model looks something like this: https://www.justice.gov/opa/pr/former-twitter-employee-found...

Re: Tainting the CSAM client-side scanning database

#212
post #36

I think it's pretty clear this is not about "CSAM", we have to stop using the term. It's just censorship, plain and simple. Client side means you'll pay from your own pocket for this wrongthing detector to work. It can even be automated so as soon as the detector gets triggered by anything, you'll get locked out of your bank accounts, until further notice I guess. If this thing gets a serious discussion in a parliame…

But parliament wants to seed your camera with mugshots of the FBI's top-ten most wanted list so the instant a false positive appears (directly on the camera, potentially even prior to writing the image to disk, potentially even prior to pressing the snapshot button)... they beacon an alert (or exfiltrate piggybacking via Bluetooth/AirTag/Covid exposure tracking mrchanisms), and bob's you're uncle.

Sounds like there's going to be a market for vintage cell phones built before that happens. Although eventually they'll stop being useful as gWhatever rolls out and telecom providers stop supporting 4/5g.

Re: Tainting the CSAM client-side scanning database

#213

Earlier quoted context omitted.

> Are they guilty of a crime? Unless there is a very specific "attempt to acquire CSAM" law then no they're not fucking guilty of any crime. If you live in a state where marijuana is illegal and you smoke some oregano because you thought it was marijuana you're not guilty of actually possessing marijuana. A criminal law is composed of a number of individual statutes. When a state is trying to prosecute someone for a…

> If a cop sells you oregano and you think it's marijuana you might have the intent to buy marijuana but there's no actual criminal act because oregano isn't illegal. If you make a law that only requires intent then congratulations, you've created thought crimes. You're a lawyer, I take it? I'm not a lawyer, and I admit your analysis of this scenario confuses me. Is there no legal difference between merely having int…

Definitely not a lawyer. You couldn't charge anyone with criminal conspiracy from the parent commenter's perspective.

Re: Tainting the CSAM client-side scanning database

#214

The article considers "an entity that is allowed to propose new entries to the CSAM database". You don't even need this! You could target a whole "social cluster" of people without having any special privileges within this system. As an example, lets say you want to attack environmental protesters. For image A, you create a meme about climate change. For image B, you procure something that looks, to humans, like CSAM…

> and if the platform fulfills its obligations, that report will bubble up to the relevant authorities, who will review it and add its fingerprint to the database. I don't think NCMEC adds random images they find to the A1 list without knowing their origin. > until the average meme-savvy environmental protester's device gets flagged for further scrutiny. Who is this an attack on? A moderation contractor maybe, but no…

Image B' *is* CSAM. A competent authority will add it when it comes to their attention.

Re: Tainting the CSAM client-side scanning database

#215

The article considers "an entity that is allowed to propose new entries to the CSAM database". You don't even need this! You could target a whole "social cluster" of people without having any special privileges within this system. As an example, lets say you want to attack environmental protesters. For image A, you create a meme about climate change. For image B, you procure something that looks, to humans, like CSAM…

> and if the platform fulfills its obligations, that report will bubble up to the relevant authorities, who will review it and add its fingerprint to the database. I don't think NCMEC adds random images they find to the A1 list without knowing their origin. > until the average meme-savvy environmental protester's device gets flagged for further scrutiny. Who is this an attack on? A moderation contractor maybe, but no…

Look at some of the cases in the news--the tech giants are more interested in ensuring that bad guys don't use their systems than in justice.

Remember that case not too long ago with a telehealth appointment, they sent a picture of something on their 2? year old's penis to the doc, asking if it was an issue. The police cleared him, but he's forever guilty in Google's eyes.

Re: Tainting the CSAM client-side scanning database

#216

The article considers "an entity that is allowed to propose new entries to the CSAM database". You don't even need this! You could target a whole "social cluster" of people without having any special privileges within this system. As an example, lets say you want to attack environmental protesters. For image A, you create a meme about climate change. For image B, you procure something that looks, to humans, like CSAM…

> Apple's proposed device scanning system had a threshold before your device would be flagged Which always struck me as odd as it would extremely easy to spin this as “Apple detected CSAM on this device but their policy is only to alert authorities once a set quantity of CSAM is found…”

It means they recognize there can be false positives. And that you can be the innocent recipient of it.

Re: Tainting the CSAM client-side scanning database

#217

Earlier quoted context omitted.

What I'm concerned about is a system that flags me for a crime based on a database I can't audit based on mechanisms with an entirely too high false positive rate. Because the database can't be audited by anyone but a select group we have to trust that it only contains actual bad images. I do not trust that such databases don't also contain images that are embarrassing to powerful/connected people. I also do not trus…

Nothing in this post (either the system implementation details or the supposed consequences of detection) is how it actually works in reality, so I don't think you should spent much time being upset that something you just made up has a security hole in it. e.g. there isn't a high false positive rate, the attack in this article doesn't work because the attacker doesn't have access to all the hashing algorithms used,…

Um, the whole thing definitely is unaudited.

> the attack in this article doesn't work because the attacker doesn't have access to all the hashing algorithms used

As far as I know, there are only two hashing algorithms used: ContentID and "the Facebook one", whose name I don't remember offhand at the moment. ContentID has leaked, been reverse engineered, been published, and been broken. A method to generate collisions in it has been published. The Facebook one has never been secret, and essentially the same method can generate collisions in it. And most users just use ContentID. [On edit: Oh, yeah, Apple has their "Neuralhash" thing.]

Are there others? If there are, I predict they're going to leak just like ContentID did, if they haven't already. You can't keep something like that secret. To actually use such an algorithm, you end up having to distribute code for it to too many places. [On the same edit: that applies to Neuralhash].

I assume you're right that the false positive rate is very low at the moment. Given the way they're done, I don't see how those hashes would match closely by accident. But the whole point of this discussion is that people have now figured out how to raise the false positive rate at will. It's a matter of when, not if, somebody finds a reason to drive the false positive rate way up. Even if that reason is pure lulz.

None of this "texts the police", but it does alert service providers who may delete files, lock accounts, or flag people for further surveillance and heightened suspicion. Much of that is entirely automated. And a lot of the other things you'd use as input, if you were writing a program to decide how suspicious you were, are even more prone to manipulation and false positives.

I believe the service providers also send many of the hits to varous "national clearinghouses", which are supposed to validate them. Those clearinghouses usually aren't the police, but they're close.

But the clearinghouses and the police aren't the main problem the false positives will cause. The main problem is the number of people and topics that disappear from the Internet because of risk scores that end up "too high".

Re: Tainting the CSAM client-side scanning database

#218
post #169

Earlier quoted context omitted.

I wish that willful false copyright claims carried the same $250,000 penalty that copyright infringement does.

That would be ridiculous. But there are penalties for false claims. There is a fine for claiming copyright you don't own, and if you go further and ask for takedowns, you are also liable for damage. The problem is that these are rarely enforced. Even a $100 fine for a false claim on YouTube would be enough to weed out bots and click farms. And for the most serious cases, have the infringer pay damage and a bigger fin…

The penalty for a *knowingly* false allegation should be the same as the penalty for the actual act.

(And, yes, I would apply that to the criminal justice world.)

Re: Tainting the CSAM client-side scanning database

#219

The article considers "an entity that is allowed to propose new entries to the CSAM database". You don't even need this! You could target a whole "social cluster" of people without having any special privileges within this system. As an example, lets say you want to attack environmental protesters. For image A, you create a meme about climate change. For image B, you procure something that looks, to humans, like CSAM…

> and if the platform fulfills its obligations, that report will bubble up to the relevant authorities, who will review it and add its fingerprint to the database. I don't think NCMEC adds random images they find to the A1 list without knowing their origin. > until the average meme-savvy environmental protester's device gets flagged for further scrutiny. Who is this an attack on? A moderation contractor maybe, but no…

It means your account gets a higher risk score, which may mean it gets given a "timeout", gets downranked in "the algorithm", gets outright shadowbanned, or may even get completely shut off. All in a completely automated way.

Re: Tainting the CSAM client-side scanning database

#220
post #164

I am against the idea of scanning for the reason that the author pointed out: It's trivial to repurpose the technology to use it in dystopian ways. I however have precisely zero concerns about impersonating hashes: 1. It's trivial to deal with tainting the database: both secondary hashing and more invasive hashes deal with that problem. 2. It's trivial to deal with impersonated hashes, all positives can be scanned on…

> It's trivial to deal with tainting the database: both secondary hashing and more invasive hashes deal with that problem. Sorry, what does "more invasive hashes" mean? > It's trivial to deal with impersonated hashes, all positives can be scanned on device in a second round with a different hash method. The double hashing would definitely help. I doubt anybody's gotten a close collision for more than one hash at a ti…

I've never played with fingerprinting per se, but long ago I did write a program to detect near-duplicate images. It was based on comparing square of the differences of extremely downscaled images. Quite good at it's intended purpose, exposed the fact that a few images were photoshops and revealed that images that were inherently low-contrast could easily be confused by that approach--beaches looked an awful lot like beaches.
Post reply on HN