Live data from Hacker News

ImageNet contains naturally occurring Apple NeuralHash collisions

blog.roboflow.com

241–250 of 530 posts

Re: ImageNet contains naturally occurring Apple NeuralHash collisions

#241

Earlier quoted context omitted.

Seems to be a misunderstanding between what the law appears to say and what the actual practice is. Law enforcement's interest is not served by trying to prosecute moderators or companies acting in good faith because they have CSAM in their possession.

There is a difference between moderators manually identifying illegal content in a stream of mostly-legal material and a process where content which has already been matched against the database and classified as almost-certainly-illegal is subjected to further internal review.

The chance of a match being CSAM is not almost certain, though. Further, Apple only gets a low-resolution version of the image. In any case, presumably such issues have been addressed, as neither the FBI nor NCMEC have raised a stink about it.

Re: ImageNet contains naturally occurring Apple NeuralHash collisions

#242

Keep in mind that Apple's claimed false positive rate (one in a trillion chance of an account being flagged innocently), and the collision rate determined by Dwyer in the article, are both derived without any adversarial assumptions. Given that NeuralHash collider and similar tools already exist, the false positive rate is expected to be much much higher. Imagine that you play a game of craps against an online casino…

Why would anyone bother with such an attack? The end result is that some peon at Apple has to look at the images and mark them as not CSAM. You've cost someone a bit of privacy, but that's it.

> "The end result is that some peon at Apple has to look at the images and mark them as not CSAM. You've cost someone a bit of privacy, but that's it."

This can be abused to spam Apple's manual review process, grinding it down to a halt. You've cost Apple time and money by making them review each such fake report.

Re: ImageNet contains naturally occurring Apple NeuralHash collisions

#243

Earlier quoted context omitted.

This isn’t true. The db is blinded. We have no way of knowing what’s in it. It would be trivial to have the same payload on each device, and extract different answers using the matching server side db which varies by country. Perhaps not trivial, but just short.

I feel that it's the kind of scheme that requires too much cooperation from too many people and organizations with conflicting incentives. It's possible some countries would not want the hashes from certain other governments in the db at all . And then what? I may be wrong, but I also believe we can know how many hashes are in the db, which means that if it contains extra hashes from dozens of governments, it would b…

Would they rather deal with the CCP shutting off iPhone sales in China? History has shown that the CCP is willing to do that if it comes down to it. (I remind you that at one time, Google was a primary search engine in China.)

Re: ImageNet contains naturally occurring Apple NeuralHash collisions

#244

Earlier quoted context omitted.

They don’t matter because if two images don’t look the same, but collide - then human processes will absolve you. This isn’t some AI that sends you straight to prison lol

Imagine this scenario. - You receive some naughty (legal!) images of a naked young adult while flirting online and save them to your camera roll. - These images have been made to collide [1] with "well known" CSAM images obtained from the dark underbelly of the internet, on the assumption that their hashes will be contained in the encrypted database. - Apple's manual review kicks in because you have enough such image…

The problem is that Apple cannot actually see the image at its original resolution because of the supposed liability of harboring CSAM, but being able to retrieve the original image would mean being able to know the complete contents of the rest of its data. To me, it sounds like Apple is trying to make a compromise between having as little knowledge of data on the server as possible and remaining in compliance with the law, but that compromise is impractical to execute.

The law states that if you find an image that's believed to be CSAM, you must report it to the authorities. If Apple's model detects CSAM on the device, sending the whole image to the moderation system for false positives carries the risk of breaking the law, because the images are likely going to be known CSAM, since that's what the database is intended to detect, so they'd be accused of storing it knowingly. Perhaps that's why the thumbnail system is needed.

So why wouldn't Apple store the files unencrypted and scan them when they arrive? That would mean Apple would remove themselves from liability by preventing themselves from gaining knowledge of which images are CSAM or not until they're scanned for, but could still send the original copy of the image with a far lower chance of false positives when something is found. That knowledge or the lack of it about the nature of the image is the crucial factor, and once they believe an image is CSAM they cannot ignore it or stop believing it's CSAM later.

That question may hold the answer to why Apple attempted to innovate in how it scans for child abuse material, perhaps to a fault.

Re: ImageNet contains naturally occurring Apple NeuralHash collisions

#247

Earlier quoted context omitted.

The answers are probably the same. What different is how hard it is to discover.

It apparently wasn't hard to "discover" the fact that this CSAM database can and will change over time. In fact, Apple explained this in detail as well as how they are attempting to avoid the problem of governments abusing the system. Are you suggesting that a different software update might be even easier to discover?

[deleted]

Re: ImageNet contains naturally occurring Apple NeuralHash collisions

#248

It doesn’t matter if there are collisions if the two images don’t actually look the same. Do people honestly believe a single CSAM flag from an “innocent” image is going to result in someone going to prison in America? PhotoDNA has existed for over a decade doing the same thing with no instances that I have heard of. If some corrupt government wants to get you they don’t need this. They can just unilaterally say you’…

We are talking about how everyone who gave Apple money now has a potential probable cause vector that they didn't before. Everyone running the software is a suspect by default. Ask black Americans how they feel about setting the bar low for probable cause. "Following the 2004 Madrid train bombings, fingerprints on a bag containing detonating devices were found by Spanish authorities. The Spanish National Police share…

> '100% verified'

Just reading those words is rage-inducing, but I'm grateful to have learnt this example of government lying. I feel like it should become an expression that societies teach to their children to warn them about abuses of power. Other mottoes synonymous with government deception and corruption come to mind, but at the risk of being too controversial I will share only their initials and dates: "SAARH" (2013), "MA" (2003), "IANAC" (1973), "NHDAEMZE" (1961), "AMF" (1940).

Re: ImageNet contains naturally occurring Apple NeuralHash collisions

#249

Earlier quoted context omitted.

What do you mean blinded? It’s already been promised there will be a way to verify the hash db on your own device.

It’s blinded in the cryptographic sense. It’s a specific term. I would go into detail, but . Suffice to say, unless you provide proof, I am reasonably confident there’s no way to verify the hash db doesn’t contain extra hashes other than the CSAM hashes provided by the US government. But I’ve been wrong many times before.

Well first of all, it's not provided by the US government. It's a non-profit, and Apple has already said they're going to look for another db from another nation and only included hashes that are the union of the two to prevent exactly this kind of attack.

If what you mean by blinded is that you don't know what the source image is for the hash, that's true. Otherwise Apple would just be putting a database of child porn on everyone's phones. You gotta find some kind of balance here.

What do you mean you can't verify it doesn't contain extra hashes? Meaning that Apple will say here are the hashes in your phone, but secretly will have extra hashes they're not telling you about? Not only is this the kind of thing that security researchers will quickly find, you're assuming a very sinister set of features from Apple that they'll only tell you half the story. If that were the case, then why offer the hashes at all? It's an extremely cynical take.

The reality is all of the complaints about this system went from this specific implementation, and then as details get revealed, it's now all about the future hypothetical situations. I'm personally concerned about future regulations, but those regulations could/would exist independently of this specific system. Further, Dropbox, Facebook, Microsoft, Google, etc all have user data unencrypted on their servers and are also just as vulnerable to said legislation. If the argument is this is searching your device, well the current implementation is its only searching what would be uploaded to a server instead. If you suggest that could change to anything on your device due to legislation, wouldn't that happen anyway? And then what is Google going to do... not follow the same laws? Both companies would have to implement new architectures and systems for complying.

I'm generally concerned about the future of privacy, but I think people (including myself initially) have gone too far in losing their minds.

Re: ImageNet contains naturally occurring Apple NeuralHash collisions

#250

> it's not obvious how we can trust that a rogue actor (like a foreign government) couldn't add non-CSAM hashes to the list to root out human rights advocates or political rivals. Apple has tried to mitigate this by requiring two countries to agree to add a file to the list, but the process for this seems opaque and ripe for abuse. If the CCP says "put these hashes in your database or we will halt all iPhone sales in…

Or why not "Hey Vietnam, Pakistan, Russia, etc put these hashes into your database please and thanks." I mean the CCP has allies that are also authoritarian. Why would they have to threaten Apple directly? This is also how you get past the Apple human verification. Just pay those Apple workers to click confirm.

> Why would they have to threaten Apple directly?

They'd do it directly because it's expedient and useful. If you're operating such a sprawling authoritarian regime, it's important to occasionally make a show of your power and control, lest anyone forget. The CCP isn't afraid of Apple, Apple is afraid of the CCP. Lately the CCP has been on a rather showy demonstration of its total control. If you're them it's useful to remind Apple from time to time that they're basically a guest in China and can be removed at any time. You don't want them to forget, you want to be confrontational with Apple at times, you want to see their acknowledged subservience; you're not looking to avoid that power confrontation, the confrontation is part of the point.

And the threat generally isn't made, it's understood. The CCP doesn't have to threaten in most cases, Apple will understand ahead of time. What gets made initially is a dictate (do this), not the threat. If something unusual happens, such as with Didi's listing on the NYSE against Beijing's wishes (whereas ByteDance did the opposite and kowtowed, pulling their plans to IPO), then, given that Didi obviously understood the confrontation risk ahead of time and tested you anyway, then you punish them. If that still isn't enough, you take them apart (or in the case of Apple, remove them from the country).

Post reply on HN