Live data from Hacker News

A catalog of naturally occurring images whose Apple NeuralHash is identical

github.com

281–290 of 304 posts

Re: A catalog of naturally occurring images whose Apple NeuralHash is identical

#281

Earlier quoted context omitted.

Don’t upload an image anywhere, else it can be reviewed.

The whole point of Apple's system is that I don't need to upload an image anywhere. Images from my phone can be stolen and reviewed with no due process, based on proprietary Apple technology.

The system as described only submits its safety vouchers when photos are uploaded to iCloud.

Not saying it will stay that way, but there are three distinct realms of objection to this system, and it's probably useful to separate them:

1. Objections that in the future, something different will happen with the technology, system, or companies; so that even if the system is unobjectionable now, we should object because of what it might be used for in the future; or how it might change. 2. Objections that Apple can't be trusted to do what they say they are doing, so that even if they say they will only refer cases after careful manual review, or that they will submit images for review that were not uploaded to iCloud, we can't believe them, so we should object. 3. Objections that hold for the system as designed and promised; in other words, even if all the actors do what they say they are doing in good faith and this monitoring never expands, it's still bad.

People who have the third kind of objection need to deal with the fact that Apple is basically putting in a system with more careful safeguards than are already in place in many Internet services, even for their "private" media storage or exchange. You likely don't know how the services you use are scanning for CSAM but if the service is at all sizeable (chat, mail, cloud storage) it's likely using PhotoDNA or something similar.

I think there are valid objections on all three bases. But there's a difference in saying "this is bad because of something that might happen" and "this is bad because of what is actually happening".

Re: A catalog of naturally occurring images whose Apple NeuralHash is identical

#282
post #98

Earlier quoted context omitted.

Since the risk of that is 1 in a trillion, a lot of people are quite happy to take that risk.

1 in a trillion with a billion devices with 10k pics is not a small chance. But what do we know? It's not like Apple is communicating any numbers anywhere so we can make a reasonable guess as to the number of false positives, and as to whether we may think it is actually worth it.

It’s not per photo, it’s per photo library. But it’s per year, so on average once every 1000 years there will be a false positive for someone.

They are communicating numbers. For example, they tested with 100 million photos and got 3 false positives. They also tested with 100 k normal porn photos and got 0 false positives.

Re: A catalog of naturally occurring images whose Apple NeuralHash is identical

#283
post #226

Earlier quoted context omitted.

So we must compare those statistics for unsolved mysteries to find the truth and correlation.

For the best criminals, you wouldn't even know that there's an unsolved mystery to be solved.

Is there even a problem in such a scenario?

Re: A catalog of naturally occurring images whose Apple NeuralHash is identical

#284
post #257
post #111

Earlier quoted context omitted.

So, the presumed attack (not against individuals, but to defeat the system) is 1. Identify some innocuous pictures that many many people have (memes, Beyoncé, whatever). 2. Produce CSAM. 3. Mangle it such that it is still CSAM visually, but NeuralHash-collides with the innocuous pictures from step 1. 4. Distribute. 5. Wait until they are (via some other mechanism) a) identified as CSAM, b) added to the NCMEC database…

> it is predicated on the assumption that you can easily mangle pictures to NeuralHash-collide with a desired target picture (out of a set of widely circulating innocuous pictures) without deteriorating the visual content too much. You can. Here is an example I created (with links to more): https://github.com/AsuharietYgvar/AppleNeuralHash2ONNX/issue... I'm so tired of people suggesting that you can't. Please explain…

> Please explain to me why you posted suggesting otherwise.

a) I didn't suggest otherwise, I said that it is predicated on that assumption, about which I was undecided, largely because b) I didn't know better.

I read that thread 7 days ago, when the collisions were a gray blob or a clearly modified dog (to Lena) or clearly modified Lena (to dog). I hadn't re-read the thread in the last 4 to 5 days, when you demonstrated the natural looking collisions (second-preimage images).

Very impressive work I wasn't aware of.

Re: A catalog of naturally occurring images whose Apple NeuralHash is identical

#285
post #284
post #257

Earlier quoted context omitted.

> it is predicated on the assumption that you can easily mangle pictures to NeuralHash-collide with a desired target picture (out of a set of widely circulating innocuous pictures) without deteriorating the visual content too much. You can. Here is an example I created (with links to more): https://github.com/AsuharietYgvar/AppleNeuralHash2ONNX/issue... I'm so tired of people suggesting that you can't. Please explain…

> Please explain to me why you posted suggesting otherwise. a) I didn't suggest otherwise, I said that it is predicated on that assumption, about which I was undecided, largely because b) I didn't know better. I read that thread 7 days ago, when the collisions were a gray blob or a clearly modified dog (to Lena) or clearly modified Lena (to dog). I hadn't re-read the thread in the last 4 to 5 days, when you demonstra…

Great answer. Thank you!

Re: A catalog of naturally occurring images whose Apple NeuralHash is identical

#286
post #29

Earlier quoted context omitted.

1. You can at least somewhat audit the software running on an iPhone, for example by means of reverse engineering. You can’t audit the server side. 2. It’s one thing to rely on proprietary services like Find My or Siri. It’s another thing to rely on a secret server-side app that has the power to destroy your life.

What I somehow fail to grasp in the first argument is that this whole system is designed specifically so that it runs client-side. AFAIK all the alternatives (as in « cloud photo services ») have been doing the exact same thing on the server side for decades. If you upload your photos to the cloud, a lot of service actually already have the power to destroy your life.

This is all about Apple users waking up to the fact that they've been had

Re: A catalog of naturally occurring images whose Apple NeuralHash is identical

#287
post #89

Earlier quoted context omitted.

Yes, but cryptographic hashes are irrelevant here because they'd allow to easily bypass CSAM by modifying/appending a single byte.

Apple is claiming to have a visual equivalent to a cryptographic hash -- one which won't change with a single byte, but only if the image is substantially different. At least their security analysis relies on that. From their whitepaper: "The threshold is selected to provide an extremely low (1 in 1 trillion) probability of incorrectly flagging a given account" If your claim is that their hash algorithm isn't cryptog…

Their security analysis is obviously incorrect.

"equivalent to a cryptographic hash" "change... only if the image is substantially different"

Both are not true, cannot be true.

Re: A catalog of naturally occurring images whose Apple NeuralHash is identical

#288

Earlier quoted context omitted.

1 in a trillion with a billion devices with 10k pics is not a small chance. But what do we know? It's not like Apple is communicating any numbers anywhere so we can make a reasonable guess as to the number of false positives, and as to whether we may think it is actually worth it.

It’s not per photo, it’s per photo library. But it’s per year, so on average once every 1000 years there will be a false positive for someone. They are communicating numbers. For example, they tested with 100 million photos and got 3 false positives. They also tested with 100 k normal porn photos and got 0 false positives.

I think their numbers are completely irrelevant, now that we know the visual hash can be gamed. Since it can be, it will be.

Basically, that 1 in a trillion number has an implicit "assuming people aren't cheating", as most mathematical models do. But it's already evident people can cheat this system.

I don't know what the odds will end up being, 1 in a trillion or 1 in 100, but they will not be based on statistical analysis. The odds will be based on cultural and social factors... how quickly do Apple reviewers get overwhelmed? How easily can script kiddies use the tools to fake hashes? Are there consequences for false reports? How many people want to get you in trouble?

Re: A catalog of naturally occurring images whose Apple NeuralHash is identical

#289

Earlier quoted context omitted.

There is an interesting constitutional quirk which arises from the scanning being done client side, specifically for US citizens. If the US Government forced Apple to add other entries to the hash table, this would constitute a warrantless Government search of the private physical property of US citizens. This is a clear-cut, unambiguous breach of the 4th Amendment. Whereas if the CSAM scanning was performed exclusiv…

Apple could also encrypt every upload to iCloud, and not have any scanning on the client, and still be able to say to the government "sure, you can have the files; we can't read them and neither can you". Apple wants to reduce your privacy from the government above and beyond what the law requires. The questions is: why?

The answer may be related to Pegasus software.

Timeline:

1. Leaked documents show Pegasus software exploits all iphones using an exploit in iMessage 2. Apple releases security update (doesn’t patch imessage. exploit) 3. Apple announces CSAM client scanning coming soon 4. Apple releases another security update (still leaves iMessage exploit unpatched and used by Pegasus)

….

Perhaps Apple is under pressure to provide a back door prior to patching a tool that may be widely used by governments around the world.

Re: A catalog of naturally occurring images whose Apple NeuralHash is identical

#290

Earlier quoted context omitted.

It's not 1 in a trillion, it's MUCH higher. We are not speaking about a situation where not a "arbitrary" picture is miss-classifieds. We are speaking about a situation where a innocent picture involving a naked or not fully clothed child is deemed similar to a non innocent picture of a naked or not fully clothed child. Now you might argue that there should not be a picture a a naked or not fully clothed child of any…

> Now you might argue that there should not be a picture a a naked or not fully clothed child of any form ever on any phone No I'm definitely not arguing that. I'm not American, where I live you'll sometimes see nude bathers in the city centre, and most definitely nude children on the beach. 1 in a trillion is derived from a dataset of 100 million photos, presumably a representable proportion of these were "similar i…

Given the problematic aspects of a "globally" available datasets and given that in many countries people are more stuck up about such thinks I highly doubt that this dataset contained a "representative proportion" given that you are more open about such thinks and take picture (in general).

I.e. there can be a massive difference between a probability "over all humans" and a probability "over people of a given culture" as long as either the given culture is in a minority or underrepresented in given data.

Given that people normally don't (knowingly) give out their private family photos when they know they culture is seen as "bad" by some people and this picture might be abused I think we can at least assume such culture(s) are underrepresented.

Through we can't say how much that changes the probability.

Post reply on HN