Live data from Hacker News

The Problem with Perceptual Hashes

rentafounder.com

61–70 of 440 posts

Re: The Problem with Perceptual Hashes

#61
Regarding false positives re:Apple, the Ars Technica article claims

> Apple offers technical details, claims 1-in-1 trillion chance of false positives.

There are two ways to read this, but I'm assuming it means, for each scan, there is a 1-in-1 trillion chance of a false positive.

Apple has over 1 billion devices. Assuming ten scans per device per day, you would reach one trillion scans in ~100 days. Okay, but not all the devices will be on the latest iOS, not all are active, etc, etc. But this is all under the assumption those numbers are accurate. I imagine reality will be much worse. And I don't think the police will be very understanding. Maybe you will get off, but you'll be in a huge debt from your legal defense. Or maybe, you'll be in jail, because the police threw the book at you.

Re: The Problem with Perceptual Hashes

#62

It really all comes down to if Apple has and is willing to maintain the effort of human evaluations prior to taking action on the potentially false positives: > According to Apple, a low number of positives (false or not) will not trigger an account to be flagged. But again, at these numbers, I believe you will still get too many situations where an account has multiple photos triggered as a false positive. (Apple sa…

Then law enforcement, a prosecutor and a jury would get involved. Hopefully law enforcement would be the first and final stage if it was merely the case that a person pressed “ok” by accident.

Re: The Problem with Perceptual Hashes

#63

The other issue with these hashes is non-robustness to adversarial attacks. Simply rotating the image by a few degrees, or slightly translating/shearing it will move the hash well outside the threshold. The only way to combat this would be to use a face bounding box algorithm to somehow manually realign the image.

In my admittedly limited experience in image hashing, typically you extract some basic feature and transform the image before hashing (eg darkest corner in the upper left or look for verticals/horizontals and align). You also take multiple hashes of the images to handle various crops, black and white vs color. This increases robustness a bit but overall yea you can always transform the image in such a way to come up with a different enough hash. One thing that would be hard to catch is if you do something like a swirl and then the consumers of that content will use a plugin or something to "deswirl" the image.

There's also something like the Scale Invariant Feature Transform that would protect against all affine transformations (scale, rotate, translate, skew).

I believe one thing that's done is whenever any CP is found, the hashes of all images in the "collection" is added to the DB whether or not they actually contain abuse. So if there are any common transforms of existing images then those also now have their hashes added to the db. The idea being that a high percent of hits from even the benign hashes means the presence of the same "collection".

Re: The Problem with Perceptual Hashes

#64

I've also implemented perceptual hashing algorithms for use in the real world. Article is correct, there really is no way to eliminate false positives while still catching minor changes (say, resizing, cropping, or watermarking). I'm sure I'm not the only person with naked pictures of my wife. Do you really want a false positive to result in your intimate moments getting shared around some outsourced boiler room for…

I fully agree with you. But while scrolling to next comment, a question came to my mind: Would it really bother me if some person that does not known my name, has never met me in real life and never will is looking at my pictures without me ever knowing about it? To be honest, I'm not sure if I'd care. Because for all I know, that might be happening right now...

Re: The Problem with Perceptual Hashes

#65
post #3

I am fairly ignorant if this space. Do any of the standard methods use multiple hash functions vs just one?

Yes, I worked on such a product. Users had several hashing algorithms they could chose from, and the ability to create custom ones if they wanted.

Re: The Problem with Perceptual Hashes

#66
This article covers three methods, all of which just look for alterations of a source image to find a fast match (in fact, that's the paper referenced). It is still a "squint to see if it is similar" test. I was under the impression there were more sophisticated methods that looked for types of images, not just altered known images. Am I misunderstanding?

Re: The Problem with Perceptual Hashes

#67
post #5

I am not exactly buying the premise here, if you train a CNN on useful semantic categories then the representations they generate will be semantically meaningful (so the error shown in blog wouldn’t occur). I dislike the general idea of iCloud having back doors but I don’t think the criticism in this blog is entirely valid. Edit: it was pointed out apple doesn’t have semantically meaningful classifier so the blog pos…

Apple's description of the training process ( https://www.apple.com/child-safety/pdf/CSAM_Detection_Techni... ) sounds like they're just training it to recognize some representative perturbations, not useful semantic categories.

Ok, good point, thanks.

Re: The Problem with Perceptual Hashes

#68

Earlier quoted context omitted.

> Do you really want a false positive to result in your intimate moments getting shared around some outsourced boiler room for laughs? these people also have no incentive to find you innocent for innocent photos. If they err on the side of false-negative, they might find themselves at the wrong end of a criminal search ("why didn't you catch this"), but if they false-positive they at worse ruin a random person's life…

Even still this has to go to the FBI or other law enforcement agency, then it’s passed on to a prosecutor and finally a jury will evaluate. I have a tough time believing that false positives would slip through that many layers. That isn’t to say CASM scanning or any other type of drag net is OK. But I’m not concerned about a perceptual hash ruining someone’s life, just like I’m not concerned about a botched millimete…

>>I have a tough time believing that false positives would slip through that many layers.

I don't, not in the slightest. Back in the days when Geek Squad had to report any suspicious images found during routine computer repairs, a guy got reported to the police for having child porn, arrested, fired from his job, named in the local newspaper as a pedophile, all before the prosecutor was actually persuaded by the defense attorney to look at these "disgusting pictures".....which turned out to be his own grand children in a pool. Of course he was immediately released but not before the damage to his life was done.

>>But I’m not concerned about a perceptual hash ruining someone’s life

I'm incredibly concerned about this, I don't see how you can not be.

Re: The Problem with Perceptual Hashes

#69
post #33

Earlier quoted context omitted.

This is really a difficult problem to solve I think. However, I think most people who are prosecuted for CP distribution are hoarding it by the terabyte. It’s hard to claim that you were unaware of that. A couple of gigabytes though? Plausible. And that’s what this CSAM scanner thing is going to find on phones.

A couple gigabytes is a lot of photos... and they'd all be showing up in your camera roll. Maybe possible but stretching the bounds of plausibility.

A couple gigabytes is enough to ruin someone’s day but not a lot to surreptitiously transfer, it’s literally seconds. Just backdate them and they may very well go unnoticed.

Re: The Problem with Perceptual Hashes

#70
post #53
post #38

Earlier quoted context omitted.

Buy a subcompact camera. Never upload such photos to any cloud. Use your local NAS / external disk / your Linux laptop's encrypted hard drive. Unless you prefer to live dangerously, of course.

Consumer NAS boxes like the ones from Synology or QNAP have "we update your box at our whim" cloud software running on them and are effectively subject to the same risks, even if you try to turn off all of the cloud options. I probably wouldn't include a NAS on this list unless you built it yourself. It looks like you've updated your comment to clarify Linux laptop's encrypted hard drive, and I agree with your line o…

At least you can deny the NAS access to the WAN by blocking it on the router or not configuring the right gateway.
Post reply on HN