Live data from Hacker News

ImageNet contains naturally occurring Apple NeuralHash collisions

blog.roboflow.com

231–240 of 530 posts

Re: ImageNet contains naturally occurring Apple NeuralHash collisions

#231
post #116

Earlier quoted context omitted.

That is not a good comparison. The extra hashes would help China find out about more borderline citizens than it otherwise would have.

Have we established that a US NGO is accepting "CSAM" hashes from China or that they are cooperating with them at all? That seems unlikely and Apple hasn't yet announced plans with how they're going to scan phones in China, I mean wouldn't China just demand outright to have full scanning capabilities of anything on the phone since you don't have any protection at all from that in China?

The main announcement was Apple was getting hashes from NCMEC but they also listed ICMEC and have said "and other groups". Much like the source database for the image hashes the list of sources is opaque and covered by vague statements.

Re: ImageNet contains naturally occurring Apple NeuralHash collisions

#232

Keep in mind that Apple's claimed false positive rate (one in a trillion chance of an account being flagged innocently), and the collision rate determined by Dwyer in the article, are both derived without any adversarial assumptions. Given that NeuralHash collider and similar tools already exist, the false positive rate is expected to be much much higher. Imagine that you play a game of craps against an online casino…

Why would anyone bother with such an attack? The end result is that some peon at Apple has to look at the images and mark them as not CSAM. You've cost someone a bit of privacy, but that's it.

Okay, let's play peon. Here are three perfectly legal and work-safe thumbnails of a famous singer: https://imgur.com/a/j40fMex. The singer is underage in precisely one of the three photos. Can you decide which one?

If your account has a large number of safety vouchers that trigger a CSAM match, then Apple will gather enough fragments to reassemble a secret key X (unique to your device) which they can use to decrypt the "visual derivatives" (very low resolution thumbnails) stored in all your matched safety vouchers.

An Apple employee looks at the thumbnails derived from your photos. The only judgment call this employee gets to make is whether it can be ruled out (based on the way the thumbnail looks) that your uploaded photo is CSAM-related. As long as the thumbnail contains a person, or something that looks like the depiction of a person (especially in a vaguely violent or vaguely sexual context, e.g. with nude skin or skin with injuries) they will not be able to rule out this possibility based on the thumbnail alone. And they will not have access to anything else.

Given the ability to produce hash collisions, an adversary can easily generate photos that fail this visual inspection as well. This can be accomplished straightforwardly by using perfectly legal violent or sexual material to produce the collision (e.g. most people would not suspect foul play if they got a photo of genitals from their Tinder date). But much more sophisticated attacks [2] are also possible: since the computation of the visual derivative happens on the client, an adversary will be able to reverse engineer the precise algorithm.

While 30 matching hashes are probably not sufficient to convict somebody, they're more than sufficient to make somebody a suspect. Reasonable suspicion is enough to get a warrant, which means search and seizure, computer equipment hauled away and subjected to forensic analysis, etc. If a victim works with children, they'll be fired for sure. And if they do charge somebody, it will be in Apple's very best interest not to assist the victim in any way: that would require admitting to faults in a high profile algorithm whose mere existence was responsible for significant negative publicity. In an absurdly unlucky case, the jury may even interpret "1 in 1 trillion chance of false positive" as "way beyond reasonable doubt".

Chances are the FBI won't have the time to go after every report. But an attack may have consequences even if it never gets to the "warrant/charge/conviction" stage. E.g. if a victim ever gets a job where they need to obtain a security clearance, the Background Investigation Process will reveal their "digital footprint", almost certainly including the fact that the FBI got a CyberTipline Report about them. That will prevent them from being granted interim determination, and will probably lead to them being denied a security clearance.

(See also my FAQ from the last thread [1], and an explanation of the algorithm [3])

[1] https://news.ycombinator.com/item?id=28232625

[2] https://graphicdesign.stackexchange.com/questions/106260/ima...

[3] https://news.ycombinator.com/item?id=28231218

Re: ImageNet contains naturally occurring Apple NeuralHash collisions

#233

> it's not obvious how we can trust that a rogue actor (like a foreign government) couldn't add non-CSAM hashes to the list to root out human rights advocates or political rivals. Apple has tried to mitigate this by requiring two countries to agree to add a file to the list, but the process for this seems opaque and ripe for abuse. If the CCP says "put these hashes in your database or we will halt all iPhone sales in…

These scenarios sound rather like "the wrong side of airlock" stories[1]. Why would China go through an elaborate scheme with fake child-porn hashes, when it can already arrest these people on made-up charges, and simply tell Apple to provide the private key for their phones, so that they can read and insert whatever real/fake evidences they want? [1] I'm stealing the expression from this excellent article: https://d…

They wouldn't. They would force apple to add hashes to things that the CCP doesn't like such as winnie the pooh memes and use turn Apple's reporting system into yet another tool to locate dissidents. How would Apple know any different. Here are some hashes, they are for CSAM trust us. They built a framework where they will call the cops on you for matching a hash value. Once governments start adding values to the database they have no reasonable way of knowing what images those actually relate to. Apple themselves said they designed it so you couldn't derive the original image from the hash. They are setting themselves up to be accessory to killing political dissidents.

Re: ImageNet contains naturally occurring Apple NeuralHash collisions

#234

Earlier quoted context omitted.

The db is encrypted and uploaded to user devices. If each country gets a different db, the payload will be different in each country, which does not make sense if it's all supposed to be CSAM. So Apple would likely just say "these were mandated by the US government for US citizens," punting the ball in their court, unless they are forbidden to say so, in which case they'll say nothing, but we all know what it means.…

This isn’t true. The db is blinded. We have no way of knowing what’s in it. It would be trivial to have the same payload on each device, and extract different answers using the matching server side db which varies by country. Perhaps not trivial, but just short.

What do you mean blinded? It’s already been promised there will be a way to verify the hash db on your own device.

Re: ImageNet contains naturally occurring Apple NeuralHash collisions

#235

Earlier quoted context omitted.

Presumably Apple would be afraid that, say, the EU becomes suspicious, issues a court order to obtain the hashes, notices they cannot audit the CCP hashes, pointedly asks "what is this", becomes absolutely livid that their citizens are spied on by a country that is not them, fines Apple out the wazoo, then extradites whoever is responsible and puts them in prison. I mean, China's not the only player in this. Putting…

I think that you overestimate the EU reaction. Every few years we learn that our Europeans leaders and some citizens have been again spied by foreign powers, such as the US, and absolutely nothing ever happened.

A friend in the military told me years ago France was the number one hacker of the US gov. It goes both ways.

This may have shifted over time as China, Russia, NK, Iran increase their attacks, but it doesn’t diminish the fact that the EU is also hacking the US without repercussions.

Re: ImageNet contains naturally occurring Apple NeuralHash collisions

#236

> Apple now has over 1.5 billion users so we are talking about a large pool of users at stake which increases the likelihood of even a low probability event manifesting This is an extremely good point. If the whole system, end to end, after all safeguards (e.g. human reviewers which can also make mistakes) has a one-in-a-billion chance to ruin a user's life, then statistically, we can expect 1-2 users to have their l…

I think you misunderstand the reporting process. If the threshold is passed, Apple reports the images to NCMEC for review. NCMEC reports it to authorities. So, this would require the failure of three organizations.

Re: ImageNet contains naturally occurring Apple NeuralHash collisions

#237

Keep in mind that Apple's claimed false positive rate (one in a trillion chance of an account being flagged innocently), and the collision rate determined by Dwyer in the article, are both derived without any adversarial assumptions. Given that NeuralHash collider and similar tools already exist, the false positive rate is expected to be much much higher. Imagine that you play a game of craps against an online casino…

Why would anyone bother with such an attack? The end result is that some peon at Apple has to look at the images and mark them as not CSAM. You've cost someone a bit of privacy, but that's it.

[deleted]

Re: ImageNet contains naturally occurring Apple NeuralHash collisions

#238

> it's not obvious how we can trust that a rogue actor (like a foreign government) couldn't add non-CSAM hashes to the list to root out human rights advocates or political rivals. Apple has tried to mitigate this by requiring two countries to agree to add a file to the list, but the process for this seems opaque and ripe for abuse. If the CCP says "put these hashes in your database or we will halt all iPhone sales in…

Or why not "Hey Vietnam, Pakistan, Russia, etc put these hashes into your database please and thanks." I mean the CCP has allies that are also authoritarian. Why would they have to threaten Apple directly? This is also how you get past the Apple human verification. Just pay those Apple workers to click confirm.

Re: ImageNet contains naturally occurring Apple NeuralHash collisions

#239
post #202

Earlier quoted context omitted.

I concede that there are overlapping issues there. But if you're saying there aren't places China goes with this sort of info that's different from the US, I don't think any debate would change your mind.

So far I have no specific reason to think China goes places that are as deeply consequential and chilling than the US. What’s Assange up to these days?

> So far I have no specific reason to think China goes places that are as deeply consequential and chilling than the US.

I would be intrigued to hear about the Uyghurs thoughts on the matter, as well as those of Assange.

Re: ImageNet contains naturally occurring Apple NeuralHash collisions

#240

Earlier quoted context omitted.

This isn’t true. The db is blinded. We have no way of knowing what’s in it. It would be trivial to have the same payload on each device, and extract different answers using the matching server side db which varies by country. Perhaps not trivial, but just short.

What do you mean blinded? It’s already been promised there will be a way to verify the hash db on your own device.

It’s blinded in the cryptographic sense. It’s a specific term. I would go into detail, but .

Suffice to say, unless you provide proof, I am reasonably confident there’s no way to verify the hash db doesn’t contain extra hashes other than the CSAM hashes provided by the US government. But I’ve been wrong many times before.

Post reply on HN