Live data from Hacker News

ImageNet contains naturally occurring Apple NeuralHash collisions

blog.roboflow.com

361–370 of 530 posts

Re: ImageNet contains naturally occurring Apple NeuralHash collisions

#361

> it's not obvious how we can trust that a rogue actor (like a foreign government) couldn't add non-CSAM hashes to the list to root out human rights advocates or political rivals. Apple has tried to mitigate this by requiring two countries to agree to add a file to the list, but the process for this seems opaque and ripe for abuse. If the CCP says "put these hashes in your database or we will halt all iPhone sales in…

Related, the Indian Government (Telecom Department) bullied Apple into building an iOS feature for reporting phone calls and SMS by threatening to stop iPhone sales in India. Apple complied. https://indianexpress.com/article/technology/mobile-tabs/app...

> an iOS feature for reporting phone calls and SMS

Why doesn't the US govt follow India's govt? I've read that Americans can receive up to 4 unsolicited calls a day.

Re: ImageNet contains naturally occurring Apple NeuralHash collisions

#362
post #353

Earlier quoted context omitted.

If the CCP says “put this arbitrary software into your next iPhone software update or we will halt all iPhone sales in China,” what do you think Apple is going to do? Isn’t the answer to both questions the same?

It's a fair question, but I think the answer is no: the questions are not the same. As much as Apple wants access to the Chinese market, it would (presumably) draw a line at some point where it would (presumably) have to choose between that market and the US market, if only because the latter is both its legal domicile and the source of most of its talent. Version A: CCP wants to exploit the hash database, there are…

> very possibly Apple moves production (not just sales) out of China

As a matter of fact, I am not sure that would be possible: It might well be that no other country has the capacity (machines and labour) to churn out that many iPhones. Would be interesting to hear if anyone has insight on that. (Tim Cook presumably knows...)

Re: ImageNet contains naturally occurring Apple NeuralHash collisions

#363

> it's not obvious how we can trust that a rogue actor (like a foreign government) couldn't add non-CSAM hashes to the list to root out human rights advocates or political rivals. Apple has tried to mitigate this by requiring two countries to agree to add a file to the list, but the process for this seems opaque and ripe for abuse. If the CCP says "put these hashes in your database or we will halt all iPhone sales in…

The CCP already runs iCloud themselves in country so this is a bit irrelevant. (Though I think this kind of capitulation to authoritarian countries is wrong, personally: https://zalberico.com/essay/2020/06/13/zoom-in-china.html ) This policy really needs to be compared to the status quo (unencrypted on cloud image scanning). When you compare in transit client side hash checks that allow for on cloud encryption and on…

Saying this is no worse than the status quo isn't a good argument. The status quo is the problem.

Re: ImageNet contains naturally occurring Apple NeuralHash collisions

#364

> it's not obvious how we can trust that a rogue actor (like a foreign government) couldn't add non-CSAM hashes to the list to root out human rights advocates or political rivals. Apple has tried to mitigate this by requiring two countries to agree to add a file to the list, but the process for this seems opaque and ripe for abuse. If the CCP says "put these hashes in your database or we will halt all iPhone sales in…

Related, the Indian Government (Telecom Department) bullied Apple into building an iOS feature for reporting phone calls and SMS by threatening to stop iPhone sales in India. Apple complied. https://indianexpress.com/article/technology/mobile-tabs/app...

I think few people are making the appropriate parallel. What we’re looking at is not necessarily government overreach, but fascism.

When the hell did it become Apple’s job to do this? Apple is not a branch of law enforcement. The government needs warrants for stuff like this. We are merging corporate and government interests here. Repeat after me, Apple is not supposed to be a branch of law enforcement.

It also says a lot about us, that we are beholden to a product. We have to ditch these products.

Re: ImageNet contains naturally occurring Apple NeuralHash collisions

#365

Earlier quoted context omitted.

There is a difference between moderators manually identifying illegal content in a stream of mostly-legal material and a process where content which has already been matched against the database and classified as almost-certainly-illegal is subjected to further internal review.

The chance of a match being CSAM is not almost certain, though. Further, Apple only gets a low-resolution version of the image. In any case, presumably such issues have been addressed, as neither the FBI nor NCMEC have raised a stink about it.

> The chance of a match being CSAM is not almost certain, though.

Not according to Apple. They're publicly claiming a one-in-a-trillion false positive rate from the automated matching. Either that's blatant false advertising or they're putting known (as in: more likely than not) CSAM in front of their human reviewers. Can't have it both ways.

> Further, Apple only gets a low-resolution version of the image.

Which makes zero difference with regard to the content being illegal. Do you think they would overlook you possessing an equally low-resolution version of the same photo?

> In any case, presumably such issues have been addressed, as neither the FBI nor NCMEC have raised a stink about it.

Selective enforcement; what else is new? It's still a huge risk for Apple to take when the ethically superior (and cheaper and simpler) solution would be to encrypt the files on the customer's device, not scanning them first, and store the backups with proper E2E encryption such that Apple has no access to or knowledge of any of the content.

Re: ImageNet contains naturally occurring Apple NeuralHash collisions

#366

Earlier quoted context omitted.

Who said anything about fake porn hashes? China can just say to Apple: We want everyone whose phones contain XYZ subversive content. Don't like it? Don't sell phones. Apple will cave. Note that you don't even have to arrest everyone. The fear is enough to prevent thoughtcrime.

Agreed. If China can force Apple to do almost anything by threatening to ban iPhone sales, why bother with fake CSAM hashes? That just adds an extra step. It's not like the Chinese government needs to take pains to trick anyone about their attitude toward "subversive" material.

This situation reminds me a lot of the controversy around Google's changes to code signing for Play Store apps. (https://news.ycombinator.com/item?id=27176690)

In both cases people are stretching to come up with hypothetical scenarios about how these systems could be abused by a government ("they could force Apple to insert non-CSAM hashes into their database" or "they could force Google to insert a backdoor into your app") while completely ignoring the elephant in the room: if a government wanted to do these things, they already have the power to do so.

If your concern is that a government might force Apple or Google to do X or pull product sales in their country, whether Apple performs on-device CSAM scanning vs scanning it on their servers, or whether Google signs your app vs you signing it doesn't materially change anything about that concern.

The outrage around this particular situation is even more confusing to me because you can opt out entirely by disabling iCloud Photos, and if you were already using iCloud Photos then the scanning was already happening on Apple's servers anyway, so the only actual change is that the scan now occurs before instead of after the upload.

Re: ImageNet contains naturally occurring Apple NeuralHash collisions

#367

Earlier quoted context omitted.

Can you explain how these theoretical political memes hash-match to an image in the NCMEC database, and then also pass the visual check? > "No, this misses the point completely. You cannot easily trigger any automated systems merely by taking photos of 17.9 year olds and sending them to people." Did I say "taking"? I am talking about sending (theoretical) actual images from the NCMEC database. This is functionally id…

Yes, I can. This is just one possible strategy: there are many others, where different things are done, and where things are done in a different order. You use the collider [1] and one of the many scaling attacks ([2] [3] [4], just the ones linked in this thread) to create an image that matches the hash of a reasonably fresh CSAM image currently circulating on the Internet, and resizes to some legal sexual or violent…

Okay, perhaps the three thumbnails was unclear. I didn't mean to illustrate any specific attack with it, just to convey the feeling of why it's difficult to tell apart legal and potentially illegal content based on thumbnails (i.e. why a reviewer would have to click "possible CSAM" even if the thumbnail looks like "vanilla" sexual or violent content that probably depicts adults). I'd splice in a sentence to clarify this, but I can't edit that particular comment anymore.

Re: ImageNet contains naturally occurring Apple NeuralHash collisions

#368

Earlier quoted context omitted.

Presumably Apple would be afraid that, say, the EU becomes suspicious, issues a court order to obtain the hashes, notices they cannot audit the CCP hashes, pointedly asks "what is this", becomes absolutely livid that their citizens are spied on by a country that is not them, fines Apple out the wazoo, then extradites whoever is responsible and puts them in prison. I mean, China's not the only player in this. Putting…

Seems pretty trivial to have different servers sides per country, and put there a different db. EU, China, Iran, US: everyone gets to spy on their own children and forbid whatever they want.

It would be good, I think, if people read Apple's threat assessment before calling it "pretty trivial":

> • Database update transparency: it must not be possible to surreptitiously change the encrypted CSAM database that’s used by the process.

> • Database and software universality: it must not be possible to target specific accounts with a different encrypted CSAM database, or with different software performing the blinded matching.

I mean, you can argue that Apple's safeguards are insufficient etc., but at least acknowledge that Apple has thought about this, outlined some solutions, and considers it a manageable threat.

ETA:

> Since no remote updates of the database are possible, and since Apple distributes the same signed operating system image to all users worldwide, it is not possible – inadvertently or through coercion – for Apple to provide targeted users with a different CSAM database. This meets our database update transparency and database universality requirements.

> Apple will publish a Knowledge Base article containing a root hash of the encrypted CSAM hash database included with each version of every Apple operating system that supports the feature. Additionally, users will be able to inspect the root hash of the encrypted database present on their device, and compare it to the expected root hash in the Knowledge Base article. That the calculation of the root hash shown to the user in Settings is accurate is subject to code inspection by security researchers like all other iOS device-side security claims.

> This approach enables third-party technical audits: an auditor can confirm that for any given root hash of the encrypted CSAM database in the Knowledge Base article or on a device, the database was generated only from an intersection of hashes from participating child safety organizations, with no additions, removals, or changes. Facilitating the audit does not require the child safety organization to provide any sensitive information like raw hashes or the source images used to generate the hashes – they must provide only a non-sensitive attestation of the full database that they sent to Apple.

[1] https://www.apple.com/child-safety/pdf/Security_Threat_Model...

Re: ImageNet contains naturally occurring Apple NeuralHash collisions

#369
post #347

Earlier quoted context omitted.

At this point, with all the easily producible collisions, the Gov't could just modify some CSAM images to match the hash of various leaked documents/etc they want to track. Then they don't even have to go thru special channels. Just submit the modified image for inclusion normally! (Not quite that simple, as they would still need to find out about the matches, but maybe that's where various NSA intercepts could help.…

Not quite, a CSAM hash match triggers another match within Apple to avoid false positives and then a human review. It wouldn't be trivial for them to extract matches out of that, and they'd only be able to track files they already know the contents for. I would think they could more easily just make your phone carrier install a malware update on your phone, rather than jumping through all of these hoops to get them a…

I tried to address the issue with finding out about the match at the end of my comment. I agree it's not exactly practical without other serious work to intercept the alerts, have 'spies' in the apple review process, etc. Much easier ways would exist at that point, but it's somewhat amusing (in a horrifying way) that some bad actor could in theory use modified CSAM as a way to detect the likely presence of non CSAM content using generated collisions.

Re: ImageNet contains naturally occurring Apple NeuralHash collisions

#370

Earlier quoted context omitted.

Can you explain how these theoretical political memes hash-match to an image in the NCMEC database, and then also pass the visual check? > "No, this misses the point completely. You cannot easily trigger any automated systems merely by taking photos of 17.9 year olds and sending them to people." Did I say "taking"? I am talking about sending (theoretical) actual images from the NCMEC database. This is functionally id…

Yes, I can. This is just one possible strategy: there are many others, where different things are done, and where things are done in a different order. You use the collider [1] and one of the many scaling attacks ([2] [3] [4], just the ones linked in this thread) to create an image that matches the hash of a reasonably fresh CSAM image currently circulating on the Internet, and resizes to some legal sexual or violent…

Your explanations are brilliant. Thank you
Post reply on HN