Live data from Hacker News

Hash collision in Apple NeuralHash model

github.com

311–320 of 725 posts

Re: Hash collision in Apple NeuralHash model

#311

Earlier quoted context omitted.

Beyond that, are we seriously going to ignore the fact that eventually this system will share the private photos of someone's naked kid with some random subcontractor? How is that even remotely OK? Or that it will find and share actual CSAM with said subcontractor?

This system does not use ML to find new CSAM images. It only checks for ones already in a known database. Your pictures of kids in the bathtub are not on the list. What is show to the reviewer is a "visual derivative" which hasn't been clearly defined. A thumbnail image? Something with a censored section? We don't really know.

Yes I'm aware that it checks against a known database but clearly there can be collisions. So eventually it will share someone's private images.

Re: Hash collision in Apple NeuralHash model

#313
post #304

Earlier quoted context omitted.

Yup! It's a legitimately crazy default. Edit: Though, to be fair, the specific hash-collision scenario would be that someone could send you something that doesn't look like CSAM and so you wouldn't reflexively delete it.

If it doesn’t look like it wouldn’t a human reviewer disregard it once it gets to that point? Personally I don’t really see the issue.

Yeah, you'd need a very specific set of processes for it to really be a problem.

Re: Hash collision in Apple NeuralHash model

#314
post #11

Apple's scheme includes operators manually verifying a low-res version of each image matching CSAM databases before any intervention. Of course, grey noise will never pass for CSAM and will fail that step. The fact that you can randomly manipulate random noise until it matches the hash of an arbitrary image is not surprising. The real challenge is generating a real image that could be mistaken for CSAM at low res + i…

Are you pinning your hopes that a false positive like this will be appropriately caught because of an army of faceless, low wage workers who stare at CSAM cases all day will immediately flag?

Apple has pretty deep pockets. Think how much that judgement is going to be when they find themselves in court for letting someone get raided over gray images.

Not that it's going to happen, since it would also require NCMEC to think the images match, but whatever. Attack me! Attack me! I want to retire.

Re: Hash collision in Apple NeuralHash model

#315
post #136

Earlier quoted context omitted.

Beyond that, are we seriously going to ignore the fact that eventually this system will share the private photos of someone's naked kid with some random subcontractor? How is that even remotely OK? Or that it will find and share actual CSAM with said subcontractor?

It only shares "visual derivatives" of images whose NeuralHash match the NeuralHash of known CSAM (either by being the same image ("perceptually") or a collision).

That to me just sounds like weasel words to avoid having to say that it shares images. Let's not beat about the bush, the "visual derivative" has to be good enough to identify what's going on in it for the manual confirmation.

Are you actually arguing in good faith here at all? Because I can't see how a "visual derivative" that's nevertheless good enough for manual confirmation is any better than the source image?

Re: Hash collision in Apple NeuralHash model

#316
post #248

Earlier quoted context omitted.

That's an incomplete statement. Currently, they must comply with warranty requests by scanning if they have the ability to scan . If they have no such ability (say, because they designed their phones from a privacy-first perspective), the law makes no requirement that they create such a capability. And that's what pisses people off about this.

Can a warrant compel them to develop the capability?

Yes. Lavabit. Those warrants demanded that Lavabit alter its system to capture passwords and/or decrypt stored email. Lavbit instead decided to stop operating and delete everything rather than comply. Such warrants have not been fully tested in courts but they do exist.

Re: Hash collision in Apple NeuralHash model

#317

How is this scenario unique to Apple but not everyone else who does scanning? e.g. Google, Facebook, Microsoft etc...

If the entity doing the scanning has a copy of the original image they can verify it is illegal before calling the police. With Apple's system they have to call the police on the basis of the image hash without verifying that anything illegal is on the phone.

You can whatsapp someone an innocent image doctored to have a hash collision with known CSAM. If they have default settings it will be saved to their photo reel, scanned by iOS and the police will be called.

Until the arresting police officer explains to them they are being arrested on suspicion of being a paedophile, they won't even know this has happened.

Re: Hash collision in Apple NeuralHash model

#318
post #106

Earlier quoted context omitted.

Don't feel bad at all. I dumped macOS entirely from production workflow. I cannot work on computer knowing that something is "scanning" me and I am glad that my "paranoid" feeling stopped me to upgrade all office macs. Billionaires at (Apple) don't give a flying f*ck about users privacy. It is all vertical integration in the name of world domination. How removed from reality they are. This is week after Pegasus/NSO a…

How likely is it that you will have enough colliding images in your photo library to even trigger a review? I'm guessing you need at least 5 images, perhaps much more, to trigger it. In any case, 1 image is definitely not enough.

Not sure why the downvotes. You are absolutely correct that a threshold in the number of positives (false or otherwise) must be met.

To be fair though we do not know what the threshold is. But I would guess even higher than 5 — I would presume 12 or more.

I'm no criminologist (IANAC) but when you read about someone getting busted with child pornography they have hundreds or thousand of images — not one, not five. They're "trading cards" for these creeps.

Re: Hash collision in Apple NeuralHash model

#319

Why is this meaningfully different than, say, what Google Photos has been doing for years? If you can get rooting malware on the target device then you could 1. Produce actual CSAM rather than a hash collision 2. Produce lots of it 3. Sync it with Google Photos This attack has been available for many years and does not need convoluted steps like hash collisions if you have the means to control somebody's phone with a…

The _initial_ implementation of client-side scanning has the same initial result, but vastly different futures. Up until now it could be (foolishly or otherwise) assumed that Apple had your personal privacy in mind. Now they have demonstrated a willingness to compromise your local device privacy. Right now it's iCloud only, but I'm sure it's a boolean configuration away somewhere to make it scan everything, regardles…

But the discussion here is about malicious actors fraudulently inserting files that trigger alarms. If they have control over your device, "content that a user has implicitly agreed to share" is not a meaningful category.

Re: Hash collision in Apple NeuralHash model

#320
post #153

Earlier quoted context omitted.

Accusations must be proven true, not untrue by the accused. And besides the "victim" in the latter case there was a whole lot of diplomatic pressure and political commotion to set him up, with carrots and sticks and the aid of friendly satellite states.

They must be proven true beyond a reasonable doubt to get a person in a funny robe to put them in jail. Normal humans are not required to prove anything in order to think them.

Which is why false rape allegations are so fucking dangerous, and are not nearly as punished as harshly as they should be.
Post reply on HN