Live data from Hacker News

A catalog of naturally occurring images whose Apple NeuralHash is identical

github.com

81–90 of 304 posts

Re: A catalog of naturally occurring images whose Apple NeuralHash is identical

#81
post #30

Earlier quoted context omitted.

The “actual argument” against the system that this provides is that Apple lied about the likelihood of hash collisions. Therefore, why trust any of their other claims?

I find this an unconvincing argument as well, you're saying that because Apple made a false claim, any claim may be valid. This is obviously not the case, what they did was to, albeit likely knowingly, calculate the hash collision probability /if each bit is a coin flip/, which comes out to pow(2, -k) for k bits. It's tiny. Of course, each bit is /not/ an independent coin flip under the NeuralHash function. So again…

> what they did was to, albeit likely knowingly, calculate the hash collision probability /if each bit is a coin flip/, which comes out to pow(2, -k) for k bits. It's tiny.

I doubt that's what they did. I think they ran tests on huge numbers of pictures, got an estimate, put in a safety factor, and determined the threshold to hit their target (and put in another safety buffer then).

Naturally occurring collisions are not going to be an issue, and adversarial ones neither, I predict. Just as with current cloud providers.

Re: A catalog of naturally occurring images whose Apple NeuralHash is identical

#82

Why are exact collisions interesting? They are not intended to be compared exactly. This algorithm doesn't even give exact matches for the same image on different hardware. https://github.com/AsuharietYgvar/AppleNeuralHash2ONNX Note: Neural hash generated here might be a few bits off from one generated on an iOS device. This is expected since different iOS devices generate slightly different hashes anyway. The reason…

The hash is 96 bits long. When hashing 1 billion pictures, that gives a collision probability of 6e-12. If it were uniformly distributed. There's no way people have hashed billions of images already. It just shows that it's pretty probably there will be collisions, and on visual inspection, it looks as if the collisions will happen on visually similar images. So if there's a naked baby pic in the CSAM database, quite a few of you 100s of child pictures can be flagged.

Re: A catalog of naturally occurring images whose Apple NeuralHash is identical

#83

Earlier quoted context omitted.

Keep in mind that Apple's claimed false positive rate (one in a trillion chance of an account being flagged innocently), and the collision rate determined by Dwyer in the blog post linked from the repo [2], are both derived without making any adversarial assumptions. Given that NeuralHash collider and similar tools already exist, the practical false positive rate is now expected to be much much higher. Imagine that y…

Yeah, of course the collision rate in an adversarial dataset is likely to be much higher. But I really wonder why you think this is an important objection, do you think a lot of people want to go to the "get flagged for child porn" casino?

The existence of a preimage attack makes Apple's system completely useless for its nominal purpose. The NeuralHash collider allows the producers and distributors of CSAM material to ensure that nearly all of the next generation of CSAM will suffer from hash collisions with perfectly innocent images.

If these new images never make it to the NCMEC database, then new CSAM content will be completely NeuralHash-proof. However, if these images eventually make their way into the NCMEC database, then everybody who has the perfectly innocent originals will be dealing with an adversarial environment.

Re: A catalog of naturally occurring images whose Apple NeuralHash is identical

#84

Earlier quoted context omitted.

2 collisions out of a million images. I'm not sure how big the CSAM database is but if it's a tens of thousands and there are millions of photos uploaded a day then Apple could have a problem on their hands. This is all extrapolating from a study that doesn't use photos representative of what people actually upload. I would suspect when most photos being uploaded are of humans the actual collision rate will be much h…

They don’t take any action unless you have 30 matches in the database, which will not happen by chance.

It could happen on purpose if I intentionally send you 30 colliding images. I don’t know how iMessage handles images, but WhatsApp for example will put them directly into your photo library (and from there directly into iCloud if you’ve got syncing enabled).

Perhaps I could even do that without revealing my motives to you.

Re: A catalog of naturally occurring images whose Apple NeuralHash is identical

#85
I still think the biggest problem is that at some point a human is going to look at a false positive, this may be picture of my naked children and this human may not have the best intentions with my picture.

That said, Nextcloud is my backend and I do not upload anything to iCloud (except for MS authenticator 2fa backups), so I'm safe right?

Re: A catalog of naturally occurring images whose Apple NeuralHash is identical

#86
post #24

Earlier quoted context omitted.

Let's not forget what the alternative is: this is about images that are uploaded on icloud anyway. The alternative is to upload the image in clear (or with ane encryption key that apple controls), and let apple run the CSAM filter on their servers. Apple now has the ability to encrypt the images before sending them to icloud, with a private key you own. Except that some percentage of images that match the CSAM finger…

They already have the decryption keys for iCloud. Undoubtedly they've already been running a similar CSAM filter server-side for ages. The only thing doing this stuff client side has done is reduce privacy and erode trust.

I don't see how the privacy is any more eroded than it already was?

Eroded trust, sure, but that is mostly because they did a terrible job communicating it.

Re: A catalog of naturally occurring images whose Apple NeuralHash is identical

#87
post #33

Apple has yet to make a valid reason for implementing client side CSAM scanning. According to Apple only images that will be uploaded to iCloud will be scanned. If this is the case there is zero reason to scan locally and you can just scan the uploaded image once it is on the server. Apple has not implemented E2E nor has it released a statement indicating this will be implemented in the future.

Most people feel that things that happen on your device are safer than things in the cloud, you have probably noticed how Apple constantly stress that this or that happens "on device".

And for the suspicious, it's of course much easier to notice if Apple would change their algorithms if they happen on device.

Re: A catalog of naturally occurring images whose Apple NeuralHash is identical

#89
post #75

I don't really get what this repository is trying to achieve and what's the point of collecting collisions. Collisions will happen, that's just how it is with hashes. It's already a public knowledge that Apple has 2 more systems (some server-side verification and a manual check later) to prevent false-positives. So what's the point of researching collisions in NeuralHash?

No. Most proper cryptographic hash systems (e.g. used for verifying files, rather than data structures) never have collisions. Try to find a SHA256 collision. Anywhere, ever, in the history of mankind. This isn't for lack of looking. A lot of very smart people have looked for them. If you find one, I bet you'll be eligible for a tenured faculty slot at a good university, if not more. A whole world of secure systems w…

Yes, but cryptographic hashes are irrelevant here because they'd allow to easily bypass CSAM by modifying/appending a single byte.

Re: A catalog of naturally occurring images whose Apple NeuralHash is identical

#90
post #41
post #18

Earlier quoted context omitted.

A catalog of a thousand pages begins with the first entry.

And a story may start with the first word, but if I present the word "Octopus" and say check out my story, you're going to be well within bounds to question me on it.

Wikipedia started as just an announcement of the idea (https://web.archive.org/web/20030414014355/http://www.nupedi... ).

How many entries do you think there were when the first live version was announced only a few hours later (http://www.nupedia.com/pipermail/nupedia-l/2001-January/0006... )?

As a different metaphor than your "Octopus", this is "first light".

"First light" in astronomy is the first time a telescope is used. It doesn't need to start with an amazing or ground-breaking image.

Post reply on HN