Live data from Hacker News

Hash collision in Apple NeuralHash model

github.com

451–460 of 725 posts

Re: Hash collision in Apple NeuralHash model

#451

How is this scenario unique to Apple but not everyone else who does scanning? e.g. Google, Facebook, Microsoft etc...

If the entity doing the scanning has a copy of the original image they can verify it is illegal before calling the police. With Apple's system they have to call the police on the basis of the image hash without verifying that anything illegal is on the phone. You can whatsapp someone an innocent image doctored to have a hash collision with known CSAM. If they have default settings it will be saved to their photo reel…

A low res copy is available to the reviewer on Apple's end to verify.

Re: Hash collision in Apple NeuralHash model

#452

Earlier quoted context omitted.

> Of course, grey noise will never pass for CSAM and will fail that step. Never? You sure that one or more human operators will never make this mistake, dooming someone's life / causing them immense pain?

I can guarantee nobody will see the inside of a courtroom, on charges of possession and distribution of child porn for possessing multiple images of grey noise (unless there is some steganography going on).

What if it is legal pornography of 21 year olds but disturbed to collide with CSAM? You are aware even defence lawyers are not allowed to look at alleged CSAM material in court right?

Re: Hash collision in Apple NeuralHash model

#453
post #152

Earlier quoted context omitted.

Not complete answers but background: apple’s system works by having your device create a hash of each image you have. The hash (a short hexadecimal string) is compared to a list of known CP image hashes, and if it matches, then your image is uploaded to Apple for further investigation. A devastating scenario for such a system is if an attacker knows how to look at a hash and generate some image that matches the hash,…

> A devastating scenario for such a system is if an attacker knows how to look at a hash and generate some image that matches the hash, allowing them to trigger false positives any time. This is my understanding too. But is this not also true for other (cloud-based) CSAM scanning systems? Why is Apple's special in this regard?

They aren't.

Apple could have saved themselves so much backlash and not have caused the outrage to be focused exclusively on them if they hadn't tried to be novel with their method of hashing, and had just announced that they were about to do exactly what all the other tech companies had already been doing for years - server side scanning.

Apple would still be accused of walking back on its claims of protecting users' privacy, but for a different reason - by trying to conform. Instead of wasting all the debate on how Apple and only Apple is violating everyone's privacy with its on-device scanning mechanism, which was without precedent, this could have been an educational experience for many people about how little privacy is valued in the cloud in general, no matter who you choose to give your data to, because there is precedent for such privacy violations that take place on the server.

Apple could have been just one of the companies in a long line of others whose data management policies would have received significant renewed attention as a result of this. Instead, everyone is focused on criticizing Apple.

There is a significant problem with people's perception of "privacy" in tech if merely moving the scan on-device causes this much backlash while those same people stayed silent during the times that Google and Facebook and the rest adopted the very same technique on the server in the past decade. Maybe if Apple had done the same, they would have been able to get away with it.

Re: Hash collision in Apple NeuralHash model

#454

Earlier quoted context omitted.

> "All user data on device and in the cloud is not scannable for CSAM or retrievable with a warrant" Who said anything about warrants? As far as I know the proposed system is proactive and requires no warrants at all. More to the point, anyone remotely sophisticated can just encrypt CSAM into a binary blob and plaster it all over the cloud providers servers. Ie, this system will possibly catch some small time pervert…

I haven’t ever said this is a good thing and that we should like it. I’m saying if the concern is that a government orders Apple to change it and do something different, then that’s a government problem and maybe we should try fixing that.

> then that’s a government problem and maybe we should try fixing that.

But why even give governments hints that people are generally OK with their devices being scanned?

We can argue about the technicalities of how abusable or resilient the current implementation is. But we can agree that it's a step towards losing privacy, yes? We didn't have scanning of iDevices before, now we do.

Because in my mind it's not a long shot to argue that once it becomes normalized that Apple can scan people's phones for CSAM when uploading to iCloud, it's just a small extra step to scan all pictures. The capability is basically in place already, it's literally removing a filter.

And then the next small step is not just CSAM but any fingerprint submitted by LE. And so it goes.

Governments can't legally compel Apple to implement this capability. But if the capability is already there, Apple can be compelled to turn over the information collected. Again, they can't do that if the capability and information doesn't exist.

E.g. if Apple can't decrypt data on your phone because they designed it such, then they can't be forced, even with a warrant, to backdoor your phone. They can legally refuse to add such capabilities.

Re: Hash collision in Apple NeuralHash model

#455

Earlier quoted context omitted.

But they won't be my private photos, the only way it gets to the point where a human see's them is that I have a bunch of child porn uploaded into iCloud, or someone puts a bunch of these grey blob images into my library. Neither of those are my images. There is no chance 30+ of my personal images have hash collisions with this database of child porn. > How society went from keeping personal photos private Also, didn…

Hash collisions dont need grey images You can take a perfectly normal image , play around with a range of its individual pixels, to get collisions too. So someone might forward you your personal pics from your meeting with them and its pixels might get edited by the messenger app you use to save the picture into your photos to specifically collide its hash with a csam image, while to you the image will look perfectly…

> So someone might forward you your personal pics from your meeting with them and its pixels might get edited by the messenger app you use to save the picture into your photos to specifically collide its hash with a csam image, while to you the image will look perfectly normal

Ok, but then that image is already not private, if I've been sending it to someone that could do that.

> This is one example , lets say you 100% trust every developer of every app that you download on apple’s phone.

I don't have to, they have to ask me before storing images in my photo library. I'm not going to give some fart generator app access to my library.

> Do you now trust every single image which might get auto downloaded to your gallery by a malicious actor , what if a random anon person messages you with 100 such colliding images , enough to cross threshold to get the authorities knocking your door

Beyond the fact that nothing writes images to my library automatically. If they sent me 100 images that are grey blobs or slightly manipulated normal images they dont get past the check anyway so no police. If they send me 100 CSAM images I'll be on the phone to the police anyway.

Re: Hash collision in Apple NeuralHash model

#456
post #372

Some people here in comments believe that whoever gonna check reported material on Apple side will never ever flag false-positive. We already know that NCMEC database itself don't exclusively contain child porn, but also some other photos that closely related ot CSAM. Even if those photos don't have actual CSAM on them. But let's ignore this fact. Do people who believe in behevolent Apple understand that CSAM don't a…

I guarantee you out of millions of alleged CSAM images hapzardly added by local police and intetnet reports, thousands are legal porn of consenting adults.

Re: Hash collision in Apple NeuralHash model

#457
post #338

Earlier quoted context omitted.

Can a warrant compel them to develop the capability?

Have they been served a warrant? Just because someone could force you to do something you don't want to do is not really a major concern if you choose, preemptively, to do the thing. The concerning part here is Apple signalling that they have executives that don't see phone scanning as a problem. That is a major black eye for their branding.

Apple's warrant canary disappeared in 2014.

Re: Hash collision in Apple NeuralHash model

#458

Earlier quoted context omitted.

Do you know how many billions of pictures of guns and drugs have ever been taken?

Sure! Lots! My Instagram account is mainly weed (legal in Canada) and Instagram censorship hits my content occasionally (because illegal cannabis selling on the platform is rampant).

And how many of those billions will Apple hash and check for?

Re: Hash collision in Apple NeuralHash model

#459

Earlier quoted context omitted.

> Apple’s method of detecting known CSAM is designed with user privacy in mind. Instead of scanning images in the cloud, the system performs on-device matching using a database of known CSAM image hashes provided by NCMEC and other child-safety organizations. Apple further transforms this database into an unreadable set of hashes, which is securely stored on users’ devices. https://www.apple.com/child-safety/pdf/CSAM…

Yeah, so how would you know something would collide with the hash?

Because the hashes are stored on the user's device?

Re: Hash collision in Apple NeuralHash model

#460
post #414

Earlier quoted context omitted.

> The warrant would properly only be to search iCloud, iCloud is encrypted, so that warrant is useless. They need to unlock and search the device.

Yes, it's encrypted, but part of this anti-CSAM strategy is a threshold encryption scheme that allows Apple to decrypt photos if a certain number of them have suspicious hashes.

Apple having any kind of ability to decrypt user contents is disconcerting. It means they can be subpoenaed for that information.
Post reply on HN