Live data from Hacker News

ImageNet contains naturally occurring Apple NeuralHash collisions

blog.roboflow.com

481–490 of 530 posts

Re: ImageNet contains naturally occurring Apple NeuralHash collisions

#481
post #449

Earlier quoted context omitted.

The objective of being mindful of the thumbnail is to fool the human reviewer responsible for alerting the police to your target's need for a good swatting - the algorithm has already flagged the image by the time it is presented as a thumbnail during review. You'd basically start off with an image known (or very likely) to be cataloged in a CP hash database. Note its NeuralHash. Find a non-CP image that would, after…

… and repeat “on the order of 30+“ times to trigger the reporting threshold, and convince them to save it into their synced-to-iCloud Photos library.

So have you never heard of catfishing? Because if you have, then you know it wouldn't be hard to do just what you described - and you're pretending otherwise for some reason.

Re: ImageNet contains naturally occurring Apple NeuralHash collisions

#482

Earlier quoted context omitted.

The objective of being mindful of the thumbnail is to fool the human reviewer responsible for alerting the police to your target's need for a good swatting - the algorithm has already flagged the image by the time it is presented as a thumbnail during review. You'd basically start off with an image known (or very likely) to be cataloged in a CP hash database. Note its NeuralHash. Find a non-CP image that would, after…

my aim was to point out that the above reverenced "image scaling attack" is easily protected against, because it is fragile to alternate scaling methods -- it breaks if you don't use the scaling algorithm the attacker planned for, and there exist secure scaling algorithms that are immune. [0] Since defeating the image scaling attack is trivial, it means that, if it is addressed, the thumbnail will always resemble the…

> However, that confusing image would almost certainly not also fool Apple's unspecified secondary server-side hashing algorithm, as referenced on page 13 of Apple's Security Threat Model Review...

Uh, on what timescale? If you mean "tomorrow" then sure, if you mean "for years" - then no. They're relying on the second perceptual hashing algorithm to remain a secret, which is insanely foolish. Just based on what I know about these CP hashlists and the laziness of programmers, I feel pretty confident that it is either an algorithm trained on the thumbnails themselves (which would be laughably bad) or it was a prior attempt that got replaced by what is now deployed on the users' hardware. Why would I think that? Because it would have been the only other thing on hand for the necessary step of generating the hash black list. So they're stuck with at least one of those forever - and will have a very limited range of potential responses to the massive infosec spotlight picking them apart... unless they want to recatalog every bit of CP all over again.

Re: ImageNet contains naturally occurring Apple NeuralHash collisions

#483
> Apple has tried to mitigate this by requiring two countries to agree to add a file to the list, but the process for this seems opaque and ripe for abuse.

They mean Russia and Belarus would need to agree to add a file to the list? Yeah, this is a very hard barrier to overcome! /s

Re: ImageNet contains naturally occurring Apple NeuralHash collisions

#484

Earlier quoted context omitted.

More broadly speaking, every part of this scheme that is currently an arbitrary Apple decision (and not a technological limitation), can easily become an arbitrary government decision. And yes, it's true that the governments could always mandate such scanning before. The difference is that it'll be much harder politically for Apple to push back against tweaks to the scheme (such as lowering the bar for manual review…

What your saying is that the government can compel Apple to develop software and include it in iOS? Haven’t they always been able to do that?

The post you responded to already addressed that exact point.

>And yes, it's true that the governments could always mandate such scanning before. The difference is that it'll be much harder politically for Apple to push back against tweaks to the scheme (such as lowering the bar for manual review / notification of authorities) if they already have it rolled out successfully and publicly argued that it's acceptable in principle, as opposed to pushing back against any kind of scanning at all.

>Once you establish that something is okay in principle, the specifics can be haggled over. I mean, just imagine this conversation in a Congressional hearing:

Re: ImageNet contains naturally occurring Apple NeuralHash collisions

#485
post #449

Earlier quoted context omitted.

The objective of being mindful of the thumbnail is to fool the human reviewer responsible for alerting the police to your target's need for a good swatting - the algorithm has already flagged the image by the time it is presented as a thumbnail during review. You'd basically start off with an image known (or very likely) to be cataloged in a CP hash database. Note its NeuralHash. Find a non-CP image that would, after…

… and repeat “on the order of 30+“ times to trigger the reporting threshold, and convince them to save it into their synced-to-iCloud Photos library.

...or just blast the messages to them on WhatsApp or any other tool that automatically dumps incoming images into the photo roll.

Re: ImageNet contains naturally occurring Apple NeuralHash collisions

#486
post #475
post #471

Earlier quoted context omitted.

This requires the attacker handling CSAM which defeats the benefit. The risk in all cases is anytime you actually handle CSAM then the attack is void since you're now actually guilty of the crime and have to do it (very few will cross that line). The point though is that this is something someone's Apple phone is doing, that their device is not. So the goal is to send a hash collided images by non-Apple channels (ema…

> the sender has committed no crime > they've done nothing illegal Perhaps that's true in the narrowest sense, but aren't the odds of generating a colliding file so low as to all but rule out coincidence and therefore strongly indicate premeditated cyber-attack (which is illegal)? If I were law enforcement, at the very least I'd want to keep tabs on these sources of false positives. Probably easy enough to convince a…

Your argument is "the technology is flawed, there let's also arrest anyone who we suspect of generating false positives".

Like security researchers. Or the people currently inspecting the algorithm. And also frankly what are you going to do about overseas adversaries? The most likely people looking at how to exploit this would explicitly be state-sponsored Russian hackers - this is right up the alley of their desire to be able to cause low level chaos without committing to a serious attack.

And at the end of the day you've still succeeded: the point is that by the time you've established it was spurious, the target has already been through the legal wringer. The legal wringer is the point.

Re: ImageNet contains naturally occurring Apple NeuralHash collisions

#487
post #472

Earlier quoted context omitted.

Why would anyone bother with such an attack? The end result is that some peon at Apple has to look at the images and mark them as not CSAM. You've cost someone a bit of privacy, but that's it.

1) be a horrible human and want to troll 2) modify close up/ambiguous adult porn to be flagged as CP with free GitHub tool 3) batch a few thousand porn photos like this to poison them 4) upload them everywhere, 4chan/Reddit/tumblr/discord/imagefap 5) some poor sap manages to save 20+ of your bait images 6) apple reviewer sees 100x100px blurry gray image of definitely porn that was flagged as CP. hits report. 7) a SWA…

This is actually a good idea. Maybe can convince apple this won't work.

Re: ImageNet contains naturally occurring Apple NeuralHash collisions

#488
As a completely unrelated side note separate to Apple's implementation.

I just read today that OnlyFans also contribute hashes to the same organisation, I wonder how many other companies have been doing the same for years.

Haven't see any uproar about this elsewhere for other services. For example, CloudFare has had this as an option on their environment for years.

Re: ImageNet contains naturally occurring Apple NeuralHash collisions

#489

Earlier quoted context omitted.

Yes, I can. This is just one possible strategy: there are many others, where different things are done, and where things are done in a different order. You use the collider [1] and one of the many scaling attacks ([2] [3] [4], just the ones linked in this thread) to create an image that matches the hash of a reasonably fresh CSAM image currently circulating on the Internet, and resizes to some legal sexual or violent…

This attack doesn’t work. If the resized image doesn’t match the CSAM image your NeuralHash mimicked, then when Apple runs it’s private perceptual hash, the hash value won’t match the expected value and it will be ignored without any human looking at it.

We have no reason to believe that Apple's second, secret perceptual hash provides any meaningful protection against such attacks. At best, we can hope that it'll allow early detection of attacks in a few cases, but chances are that's the best it can do. We might not ever learn: Apple now has a very strong incentive not to admit to any evidence of abuse or to any faults in their algorithm.

(Sorry, this is going to be long. I know understand most/all of this stuff, it's mostly there to provide a bit of context for the users reading our exchange)

The term "hash function" is a bit of a misnomer here. When people hear "hash", they tend to think about cryptographic hash functions, such as SHA256 or BLAKE3. When two messages have the same hash value, we say that they collide. Fortunately, cryptographic hash functions have several good properties associated with them: for example, there is no known way to generate a message that yields a given predetermined hash value, no known way to find two different messages with the same hash value, and no known way to make a small change to a message without changing the corresponding hash value. These properties make cryptographic hash functions secure, trustworthy and collision-resistant even in the face of powerful adversaries. Generally, when you decide to use two unrelated cryptographic hash algorithms instead of one, executing a preimage attacks against both hashes becomes much more difficult for the adversary.

However, as you know, the hash functions that Apple uses for identifying CSAM images are not "cryptographic hash functions" at all. They are "perceptual hash functions". The purpose of a perceptual hash is the exact opposite of a cryptographic hash: two images that humans see/hear/perceive (hence the term perceptual) to be the same or similar should have the same perceptual hash. There is no known perceptual hash function that remains secure and trustworthy in any sense in the face of (even unsophisticated) adversaries. In particular, preimage attacks against perceptual hashes are very easy, compared to the same attacks against cryptographic hashes.

Using two unrelated cryptographic hashes meaningfully increases resistance to collision and preimage attacks. Using ROT13 twice does not increase security in any meaningful sense. Using two perceptual hashes, while not as bad, is still much closer to the "using ROT13 twice for added security" than to the "using multiple cryptographic hashes" end.

Finding a SHA1 collision took 22 years, and there are still no effective preimage attacks against it. Creating the NeuralHash collider took a single week. More importantly, even if you were to use two unrelated perceptual hash functions, executing a preimage attacks against both hashes need not become much more difficult for the adversary: easy * easy is still easy. Layering cryptography upon cryptography is meaningful, but only as long as one of the layers is actually difficult to attack. This is not the case for perceptual hashes. In fact, in many similar contexts, these adversarial attacks tend to transfer: if they work against one technique or model, they often work against other models as well [3]. In the attack discussed above, the adversary has nearly full control over the "visual derivative", so even a very unsophisticated adversary can subject the target thumbnail itself to the collider before performing the resizing attack, and hope that it transfers against the second hash. If the second hash is a variant of NeuralHash (somewhat likely, it could even be NeuralHash performed on the thumbnail itself; we don't know anything about it!), or if it's a ML model trained on the same or similar datasets (quite likely), or if it's one of the known algorithms (say PhotoDNA) then some amount of transfer is likely to happen. And given an adversary that is going to distribute a large number of photos anyway, a 10% success rate is more than enough. Given the diminished state space (fixed size thumbnails, almost certainly smaller than 64x64 for legal reasons), a 10% success rate is completely plausible even with these naive approaches. An adversary that has some (even very little information) about the second hash algorithm can do much more sophisticated stuff, and perform much better.

But what if we boldly rule out all transfer results? Doesn't Apple keep their algorithm secret?! Can we think of the weights (coefficients) of the second perceptual hash as some kind of secret key in the cryptographical sense? Alas, no. Apple would have to make sure that all the outputs of the secret perceptual hash are kept secret as well. Due to the way perceptual hashing algorithms work, they provide a natural training gradient having access to sufficiently many input-outputs examples is probably enough to train a high-fidelity "clone" that allows one to generate adversarial examples and perform successful preimage attacks even if the weights of the clone are completely different from the secret weights of the original network. This can be done with standard black box techniques [4]. It's much harder (but nowhere near crypto hard, still perfectly plausible) to pull this off when they have access to one bit of output (match or no match). A single compromised Apple employee can gather enough data to do this given the ability to observe some inputs and outputs, even if said employee has no access to the innards or the magic numbers. The hash algorithm is kept secret because if it wasn't, an attack would be completely trivial: but an adversary does not need to learn this secret to mount an effective attack.

These are just two scenarios. There are many others. "Nobody has ever demonstrated such an attack working end-to-end" is not a good defense: it's been two weeks since the system was rolled out, and once an attack is executed, we probably won't learn about it for years to come. But the attacker can be rewarded way before "due process" kicks in: e.g. if a victim ever gets a job where they need to obtain a security clearance, the Background Investigation Process will reveal their "digital footprint", almost certainly including the fact that the NCMEC got a report about them, even if the FBI never followed up on it. That will prevent them from being granted interim determination, and will probably lead to them being denied a security clearance. If you pull off this attack on your political opponents, you can prevent them from getting government jobs, possibly without them ever learning why. And again, this is one single proposed attack. There were at least 6 different attacks proposed by regular HN users in the recent threads!

As a more general observation, cryptography tends to be resistant to attacks only if one can say things such as "the adversary cannot be successful unless they know some piece of information k, and we have very good mathematical reasons (e.g. computational hardness) to believe that they can't learn k". The technology is flawed: even the state-of-the-art in perceptual hashes does not satisfy this criterion. Currently, they are at best technicool gadgets, but layering technicool upon technicool cannot make their system more secure.And Apple's system is a high-profile target if there ever was one.

Barring a major breakthrough in perceptual hashing (one that Apple decided to keep secret and leave out of both whitepapers), the claim that the secret second hash will prevent collision attacks is not justified. The chances of such a secret breakthrough are very slim: it'd be like learning that SpaceX has already built a base on the Moon and has been doing regular supply runs with secret spaceships. Vaguely plausible in theory (SpaceX has people who do rocketry, Apple has people who do cybersecurity), but vanishingly unlikely in practice.

And that's before we mention that the mere existence of the collider made the entire exercise completely pointless: the real pedos can now use the collider to effectively anonymize their CSAM drops, making sure that all of their content collides with innocnent photos, and ensuring that none of the images will be picked up by NeuralHash anyway. For all practical purposes, Apple's CSAM detection is now _only_ an attack vector, and nothing else.

[3] https://arxiv.org/abs/1809.02861 [4] https://towardsdatascience.com/adversarial-attacks-in-machin...

Re: ImageNet contains naturally occurring Apple NeuralHash collisions

#490
post #472

Earlier quoted context omitted.

Why would anyone bother with such an attack? The end result is that some peon at Apple has to look at the images and mark them as not CSAM. You've cost someone a bit of privacy, but that's it.

1) be a horrible human and want to troll 2) modify close up/ambiguous adult porn to be flagged as CP with free GitHub tool 3) batch a few thousand porn photos like this to poison them 4) upload them everywhere, 4chan/Reddit/tumblr/discord/imagefap 5) some poor sap manages to save 20+ of your bait images 6) apple reviewer sees 100x100px blurry gray image of definitely porn that was flagged as CP. hits report. 7) a SWA…

Hasn't this ship already sailed though?

At least I'm given to understand that they already scanned all these photos uploaded to iCloud anyway (in the same way many other similar providers do). Whether it happens on the device or the server doesn't seem to make any difference to this attack.

(That's not to say that (a) the scanning of stuff on a server was a good idea in the first place or (b) encouraging politicians to use your own device to spy on you is a good idea or (c) this isn't the thin end of a very painful wedge, just that we've not opened a new vulnerability)

Post reply on HN