Live data from Hacker News

Hash collision in Apple NeuralHash model

github.com

591–600 of 725 posts

Re: Hash collision in Apple NeuralHash model

#591
post #356

Earlier quoted context omitted.

But how does this exact attack and scenario not also apply to Google, Facebook, Microsoft etc... who are also doing the same thing on their clouds servers?

I don't know what those companies do, hopefully someone who does know will chime in and answer.

They essentially do the same, just in the cloud meaning they access all images directly.

Re: Hash collision in Apple NeuralHash model

#592

Earlier quoted context omitted.

> A devastating scenario for such a system is if an attacker knows how to look at a hash and generate some image that matches the hash, allowing them to trigger false positives any time. This is my understanding too. But is this not also true for other (cloud-based) CSAM scanning systems? Why is Apple's special in this regard?

They aren't. Apple could have saved themselves so much backlash and not have caused the outrage to be focused exclusively on them if they hadn't tried to be novel with their method of hashing, and had just announced that they were about to do exactly what all the other tech companies had already been doing for years - server side scanning. Apple would still be accused of walking back on its claims of protecting users…

Perception aside, Apple’s system is somewhat better for privacy, since Apple needs to access much less data server side.

Re: Hash collision in Apple NeuralHash model

#593
post #348

Earlier quoted context omitted.

> A devastating scenario for such a system is if an attacker knows how to look at a hash and generate some image that matches the hash, allowing them to trigger false positives any time. This is my understanding too. But is this not also true for other (cloud-based) CSAM scanning systems? Why is Apple's special in this regard?

I don't know. I don't know what other systems use, I just know what I've read recently about Apple's.

This pretty much sums this entire drama up, I think.

Re: Hash collision in Apple NeuralHash model

#594
post #27

Earlier quoted context omitted.

> 7. Apple reviewer confuses a featureless blob of gray with CSAM material, several times A better collision won't be a grey blob, it'll take some photoshopped and downscaled picture of a kid and massage the least significant bits until it is a collision. https://openai.com/blog/adversarial-example-research/

So the person would have to accept and save an image that when looks enough like CSAM to confuse a reviewer…

I don't see how a perfectly legal and normal explicit photograph of someone's 20-year-old wife would be indistinguishable to an Apple reviewer from CSAM, especially since some people look much younger or much older than their chronological age. So first, there would be the horrendous breach of privacy for an Apple goon to be looking at this picture in the first place, which the person in the photograph never consented to, and second, could put the couple in legal hot water for absolutely no reason.

Re: Hash collision in Apple NeuralHash model

#595
post #70

Earlier quoted context omitted.

"How can you use it for targeted attacks?" Just insert a known CSAM image on target's device. Done. I presume this could be used against a rival political party to ruin their reputation - insert bunch of CSAM images on their devices. "Party X is revealed as an abuse ring". This goes oh-so-very-nicely with Qanon conspiracy theories which even don't require any evidence to propagate widely. Wait for Apple to find the i…

> Just insert a known CSAM image on target's device. Done. What do you mean “just”? That’s not usually very simple. It needs to go into the actual photo library. Also, you need like 30 of them inserted. > I presume this could be used against a rival political party Yes, but it’s not much different from now, since most cloud photo providers scan for this cloud-side. So that’s more an argument against scanning all toge…

WhatsApp for example automatically stores images you receive in your Photos library, so that removes a step, and those will thus be automatically uploaded to iCloud.

The one failsafe would be Apple's manual reviewers, but we haven't heard much about that process yet.

Re: Hash collision in Apple NeuralHash model

#596

Why is this meaningfully different than, say, what Google Photos has been doing for years? If you can get rooting malware on the target device then you could 1. Produce actual CSAM rather than a hash collision 2. Produce lots of it 3. Sync it with Google Photos This attack has been available for many years and does not need convoluted steps like hash collisions if you have the means to control somebody's phone with a…

The _initial_ implementation of client-side scanning has the same initial result, but vastly different futures. Up until now it could be (foolishly or otherwise) assumed that Apple had your personal privacy in mind. Now they have demonstrated a willingness to compromise your local device privacy. Right now it's iCloud only, but I'm sure it's a boolean configuration away somewhere to make it scan everything, regardles…

How have they demonstrated that they will compromise your privacy? They compromise it less than the systems scanning on the cloud, since like this most of the match is only known to the device and not the cloud.

If you don’t trust Apple to not lie, then the entire discussion becomes a bit moot, I think, since then they could pretty much do anything, with or without this system.

Re: Hash collision in Apple NeuralHash model

#597

Why is this meaningfully different than, say, what Google Photos has been doing for years? If you can get rooting malware on the target device then you could 1. Produce actual CSAM rather than a hash collision 2. Produce lots of it 3. Sync it with Google Photos This attack has been available for many years and does not need convoluted steps like hash collisions if you have the means to control somebody's phone with a…

The big difference is that cloud providers do the scanning on their own infrastructure. If you don't want something scanned, you don't upload it to the could, that simple. But here, your own device is snitching on you.

No. Only pictures being uploaded to iCloud Photo Library are being scanned.

Re: Hash collision in Apple NeuralHash model

#599

Earlier quoted context omitted.

It didn't even last a week. Introducing the people asking us to trust them with the responsibility of mass, automated crime accusations.

Why do you think somebody will accuse you of a crime because you have a photo of a grey blob? The real world isn't as stupid as the computer one. The justice system is not deterministic and automatic. Nobody is going to look at this grey blob and go "welp, we have no choice but to throw you in prison forever"

The real world and the computer world intersect. This is precisely typified by what is being discussed, surely?

Apple is trying to automate and computerise a process that was not automated previously, apply it to a huge number of people, and with disastrous potential consequences should their wonderful design be found lacking.

And within days, they have already utterly failed to provide one of their own self-stated and incredibly obvious requirements. What else could possibly go wrong?

Further, if you think that because the first pre-image failure found (in days...) was a grey blob, that this means all future possible cultivated collisions will only ever be grey blobs, well - good luck with that.

Re: Hash collision in Apple NeuralHash model

#600

Earlier quoted context omitted.

It just seems to me that there are two very different conversations happening at the same time, with people swapping back and forth between them 1. There can be false positives or other mechanisms for innocent people to get flagged. 2. It is bad to do this sort of check on the local disk. The discussion at hand started as entirely #1. But now you've swapped to #2, talking about government spying on local files. It ma…

I'd argue they are both important points that are tightly interconnected. #2 is bad on it's own and in my mind it shouldn't even get to the merits of discussing #1 because at that point the debate is already lost in favor of surveillance. But beyond that, #1 is highly problematic, doubly so given the fact that government surveillance is basically a non-decreasing function.

I’d argue that #2 is good, since it offers much more privacy than scanning in the cloud. This is based on reading and understanding the technical summary and paper linked from there.
Post reply on HN