Live data from Hacker News

A catalog of naturally occurring images whose Apple NeuralHash is identical

github.com

251–260 of 304 posts

Re: A catalog of naturally occurring images whose Apple NeuralHash is identical

#251
post #97

Earlier quoted context omitted.

If you can get exact collisions, this can be gamed. For example, suppose there are two rival gangsters. One wants to set the police on his rival. He knows that a certain (innocuous) image is on his rival's phone. So he pays someone to generate a fake child-porn image with the same neuralhash, and ensure that it gets into the child porn DB. Then, apple reports the rival to the police, and they come and investigate him…

I think before a criminal investigation, or any investigation at all is pursued, a human verifying the images would dismiss the false positive. I would think surreptitiously placing actual child porn on a rival's phone/computer would be much, much more effective. Cybercriminals could likely do all this remotely. Phish for apple account login, upload images. Done.

> a human verifying the images would dismiss the false positive

How is a human supposed to distinguish that a visual derivative (a low res sobel filtered image, presumably) of ordinary, lawful, adult pornography isn't child porn when the system has already identified it as such?

I agree that using real child porn is an attack too, but at least in that case you could say the system was doing as designed (even though what its doing shouldn't be something that we want) ... but it's not even guaranteed to do as designed.

Re: A catalog of naturally occurring images whose Apple NeuralHash is identical

#252
Is it possible for the courts to use this system to search a defendant's phone for leaked documents say? Like if NSA learns that one of a small group leaked document X, can they get a court to force Apple to add the hash of Document X to the database on that group of people's phones? If so, I bet this becomes the new norm for investigating leaks.

Re: A catalog of naturally occurring images whose Apple NeuralHash is identical

#253
post #234
post #213

Earlier quoted context omitted.

The database of hashes is controlled by the government. They can put whatever they want in there.

False. See my other reply. https://news.ycombinator.com/item?id=28303966

I don't think it's entirely false. NCMEC is not a regular non-profit. They have special clearance to do things that regular citizens and non-profits are not allowed to do.

Re: A catalog of naturally occurring images whose Apple NeuralHash is identical

#254
post #204
post #124

Earlier quoted context omitted.

Perceptual hashes are only used to reduce the search space for human review. Apple doesn’t have images in the CSAM database to do a comparison, but if it’s just a picture of a door their going to reject it. Also, because human review is an expense Apple’s incentives are to minimize the number of times it happens, thus the requirement for multiple collisions.

Apple's human review is largely useless. Trolls will be able to easily use tools slightly modify ambiguous adult porn to collide with a "known CP hash". A human reviewer will see a blurry grayscale derivative of adult pornographic content and hit "report" every time.

This is the threat model I am looking at. It is number one with a bullet. We have already had a court case where an adult actress had to show up in court and prove that she was adult when experts testified that the images were of a non-adult woman.

Baby in the sink? No. But a bunch of the aforementioned? Yeah.

Re: A catalog of naturally occurring images whose Apple NeuralHash is identical

#255
post #74

A very relevant point on this entire discourse about Apple’s on-device CSAM scanning: According to the U.S. law, key snippets of which are quoted on the Stratechery blog (by Ben Thompson), Apple isn’t obligated to scan for CSAM. It’s only obligated to act on CSAM if it finds them. While it’s good for Apple to scan on its systems (iCloud) like Facebook, Google and other companies do on their servers, it’s inappropriat…

They say only if they find 30 matching images, they'd act. So if they find 20 or 29 and don't report them, they are actually breaking the law!! I am wondering why they chose that magical number!

They think they've outsmarted the law on that one. The system is setup to generate a bunch of false positives on its own, so a "match" which is below threshold may not actually be a match-- apple can plausibly deny it.

It's unclear if their claimed threshold of 30 is before or after the false positives they intentionally introduce. I'm going to guess it's before.

Re: A catalog of naturally occurring images whose Apple NeuralHash is identical

#256
post #3

Sigh, for the last time, it doesn't actually matter if the NeuralHash is identical. You need multiple images matching, and then the images are compared by another system on Apple's end, which you don't know anything about. The system is specifically designed so that colliding images does not pose a threat to the user. NeuralHash and the CSAM scanning is grotesque, but please, criticize it for what it is, not some bul…

(A threshold of) matchings result in the private keys for the images being leaked to Apple, where they're vulnerable to:

(1) Review by apple staff (2) Access and leaking by other apple staff (3) Access by hackers who have compromised their system (4) Access by parties coercing apple/staff, including via national security letters.

All of which compromise the privacy of the user. This matters or the neuralhash comparison wouldn't exist in the first place.

Totally agree that the whole system is grotesque-- but that doesn't stop it also being grotesque in every detail as well. The fact that there are false positives when they easily could have designed a system that had none (at the expense of increased false negatives) shows that Apple doesn't especially value customer privacy even if you accept their vigilante privacy invasion. The fact that it's possible to construct adversarial false positives and that their reports didn't disclose this fact shows they either don't know what they're doing or they're not being honest about the risks (or both).

Re: A catalog of naturally occurring images whose Apple NeuralHash is identical

#257
post #111
post #44

Earlier quoted context omitted.

I think you missed the point of the first paragraph. The point is that you can now hide child porn by making its hash collide with innocent images. They won't ever make it to manual review. Ergo, NeuralHash is now useless.

So, the presumed attack (not against individuals, but to defeat the system) is 1. Identify some innocuous pictures that many many people have (memes, Beyoncé, whatever). 2. Produce CSAM. 3. Mangle it such that it is still CSAM visually, but NeuralHash-collides with the innocuous pictures from step 1. 4. Distribute. 5. Wait until they are (via some other mechanism) a) identified as CSAM, b) added to the NCMEC database…

> it is predicated on the assumption that you can easily mangle pictures to NeuralHash-collide with a desired target picture (out of a set of widely circulating innocuous pictures) without deteriorating the visual content too much.

You can. Here is an example I created (with links to more): https://github.com/AsuharietYgvar/AppleNeuralHash2ONNX/issue...

I'm so tired of people suggesting that you can't. Please explain to me why you posted suggesting otherwise.

I've contemplated making some that are also photodna matches, I expect that it's possible. But access to photodna is only through some awful windows tools, and AFAICT people would just keep posting denials even after an example was posted-- so it's not worth the effort at least not worth it just to further the public discussion.

Re: A catalog of naturally occurring images whose Apple NeuralHash is identical

#258
post #104

Earlier quoted context omitted.

"While it’s good for Apple to scan on its systems (iCloud) like Facebook, Google and other companies do on their servers, it’s inappropriate to do it on individual devices" I would even challenge the justification to do this on servers, unless the data is public. If it's behind a personal login, you might as well consider it personal property/data. I find the distinction of where data is stored not very meaningful. A…

Absolutely. If I own my data, someone processing this data on my behalf has no right or obligation to scan it for illegal content. The fact that this data sometimes sits on hard drives owned by another party just isn't a relevant factor. Presumably I still own my car when it sits in the garage at the shop. They have no right or obligation to rummage around looking for evidence of a crime. I don't see abstract data as…

Fully agree. Only this week did I learn that companies were already doing this on servers since at least 2018-2019. Can't say I ever read anything about it before.

I find it quite shocking that a foundational element of criminal justice, innocent until proven guilty and needing a reason to search individual property, is tossed aside like it's nothing.

Re: A catalog of naturally occurring images whose Apple NeuralHash is identical

#259

The technology is not why the Apple system is unwanted. It's just extra fuel for the fire. This system is unwanted because it puts a spy literally in your house and in your hands. It's bad enough that cloud everything blurs the line between what's yours and what's mine. Placing any law enforcement tech on a user's own device takes that line between "public" and "private" and completely erases it.

Absolutely. The problem is Apple introducing a spy into your home.

This alone should be bad enough, but some people are rather trusting. Showing that the spy is also tripping balls both exposes additional risks and emphasizes that Apple neither has their best interest at heart nor is putting adequate care into their actions. The latter gives people reason to question apple's claims of additional protection mechanisms that are non-falsifiable.

Re: A catalog of naturally occurring images whose Apple NeuralHash is identical

#260
post #140
post #75

I don't really get what this repository is trying to achieve and what's the point of collecting collisions. Collisions will happen, that's just how it is with hashes. It's already a public knowledge that Apple has 2 more systems (some server-side verification and a manual check later) to prevent false-positives. So what's the point of researching collisions in NeuralHash?

I'd argue that the hash collisions (both natural and synthetic) that I've seen give me more confidence in the system, not less. On the natural hash collisions (of which there are two ), we have objects of similar shape against a solid background. It seems that a natural hash collision of a CSAM image would be unlikely (or if it does occur, it would be something that perhaps is also an infringing image). As for the sy…

> this is a sh*t picture in this meme and download something else.

NO. Adversarial preimages can be created that look like perfectly normal images. Please stop repeating this falsehood.

Here are some examples I generated (with a link to more):

https://github.com/AsuharietYgvar/AppleNeuralHash2ONNX/issue...

Post reply on HN