Live data from Hacker News

The Problem with Perceptual Hashes

rentafounder.com

391–400 of 440 posts

Re: The Problem with Perceptual Hashes

#392
post #386

Earlier quoted context omitted.

If I ask you to store my images, and you therefore have access to the images, you can scan them for stuff using your computers . The scope is limited to the images I ask you to store, and your computers are doing what you ask them to. If you reprogram my computer to scan my images stored on my computer … different thing entirely. I don't have a problem with checking them for child abuse (in fact, I'd give up quite a…

So far, this is only for iCloud photos so currently it seems highly similar to what we have now except that it’s on the device and could be done with end to end encryption, unlike the current approach. For me, the big concern is how it could be expanded. This is a real and valid problem but it’s certainly not hard to imagine a government insisting it needs to be expanded to cover all photos, even for people not using…

Yes. If Apple takes the “we're not going to” stance, then this could be okay… but they've been doing that less and less, and they only ever really did that in the US / Australia. Apple just isn't trustworthy enough.

Re: The Problem with Perceptual Hashes

#393
post #285

Earlier quoted context omitted.

The article posted, as well as many others we've seen recently, demonstrate that collisions are possible, and most likely inevitable with the number of photos to be scanned for iCloud, and Apple recognizes this themselves. It doesn't necessarily mean that all flagged photos would be of explicit content, but even if it's not, is Apple telling us that we should have no expectation of privacy for any photos uploaded to…

Apple already use the same algorithm on photos in email, because email is unencrypted. Last year Apple reported 265 cases according to the NYT. Facebook reported 20.3 million. Devolving the job to the phone is a step to making things more private, not less. Apple don’t need to look at the photos on the server (and all cloud companies in the US are required to inspect photos for CSAM) if it can be done on the phone, r…

> all cloud companies in the US are required to inspect photos for CSAM)

This is extremely disingenuous. If their devices uploaded content with end to end encryption there would be no matches for CSAM.

If they were required to search your materials generally, then they would be effectively deputized-- acting on behalf of the government-- and your forth amendment protection against unlawful search would be would extended to their activity. Instead we find that the both cloud providers and the government have argued and the courts have affirmed the opposite:

In US v. Miller (2017)

> Companies like Google have business reasons to make these efforts to remove child pornography from their systems. As a Google representative noted, “[i]f our product is associated with being a haven for abusive content and conduct, users will stop using our services.” McGoff Decl., R.33-1, PageID#161.

> Did Google act under compulsion? Even if a private party does not perform a public function, the party’s action might qualify as a government act if the government “has exercised coercive power or has provided such significant encouragement, either overt or covert, that the choice must in law be deemed to be that of the” government. [...] Miller has not shown that Google’s hash-value matching falls on the “compulsion” side of this line. He cites no law that compels or encourages Google to operate its “product abuse detection system” to scan for hash-value matches. Federal law disclaims such a mandate. It says that providers need not “monitor the content of any [customer] communication” or “affirmatively search, screen, or scan” files. 18 U.S.C. § 2258A(f). Nor does Miller identify anything like the government “encouragement” that the Court found sufficient to turn a railroad’s drug and alcohol testing into “government” testing. See Skinner, 489 U.S. at 615. [...] Federal law requires “electronic communication service providers” like Google to notify NCMEC when they become aware of child pornography. 18 U.S.C. § 2258A(a). But this mandate compels providers only to report child pornography that they know of; it does not compel them to search for child pornography of which they are unaware.

Re: The Problem with Perceptual Hashes

#394
post #386

Earlier quoted context omitted.

So far, this is only for iCloud photos so currently it seems highly similar to what we have now except that it’s on the device and could be done with end to end encryption, unlike the current approach. For me, the big concern is how it could be expanded. This is a real and valid problem but it’s certainly not hard to imagine a government insisting it needs to be expanded to cover all photos, even for people not using…

Yes. If Apple takes the “we're not going to” stance, then this could be okay… but they've been doing that less and less, and they only ever really did that in the US / Australia. Apple just isn't trustworthy enough.

Also that since the system is opaque by design it’d be really hard to tell if details changes. Technically I understand why that’s the case but it makes the question of trust really hard.

Re: The Problem with Perceptual Hashes

#395
post #253

Earlier quoted context omitted.

The number of possible different images doesn't matter, it's only the number of actually different images encountered in the world. This number cannot be anywhere near 2^256, that would be physically impossible.

But you cannot know that a-priori so it’s either an attack vector for image manipulation or straight up false positives. Assume we had this perfect hash knowledge. I’d create a compression algorithm to uniquely map between images and the 256 bit hash space, which we probably agree is similarly improbable. It’s on the order of 1000x to 10000x more efficient than JPEG and isn’t even lossy.

You’re going to have to explain that—what is an attack vector for image manipulation? What is an attack vector for false positives?

> Assume we had this perfect hash knowledge.

It’s not a perfect hash. Nobody’s saying it’s a perfect hash. It’s not. It’s a perceptual hash. It is specifically designed to map similar images to similar hashes, for the “right” notion of similar.

Re: The Problem with Perceptual Hashes

#396

Earlier quoted context omitted.

Couldn't the hack just be as simple as sending someone an iMessage with the images attached? Or somehow identify/modify non-illegal images to match the perceptual hash -- since it's not a cryptographic hash.

Does iCloud automatically upload iMessage attachments?

No, iMessages are stored on the device until saved to iCloud. However, iMessages may be backed up to iCloud, if enabled.

The difference is photos saved are catalogued, while message photos are kept in their threads.

Will Apple scan photos saved via iMessage backup?

Re: The Problem with Perceptual Hashes

#398

Earlier quoted context omitted.

Does iCloud automatically upload iMessage attachments?

No, iMessages are stored on the device until saved to iCloud. However, iMessages may be backed up to iCloud, if enabled. The difference is photos saved are catalogued, while message photos are kept in their threads. Will Apple scan photos saved via iMessage backup?

I would assume yes, that this would cover iMessage backups since it is uploaded to their system.

Re: The Problem with Perceptual Hashes

#399
post #110

The problem of hash or NN based matching is, the authority can avoid explaining the mismatch. Suppose the authority want to false-arrest you. They prepare a hash that matches to an innocent image they knew the target has in his Apple product. They hand that hash to the Apple, claiming it's a hash from a child abuse image and demand privacy-invasive searching for the greater good. Then, Apple report you have a file th…

The police can arrest you for laws that don't exist but they think exist. They don't need to any of this stuff.

Re: The Problem with Perceptual Hashes

#400

Earlier quoted context omitted.

> but there is still a very strong culture of free speech, particularly political speech. Free speech didn't seem so important recently when the SJW crowd started mandating to censor certain words because they're offensive.

Free speech doesn't mean the speaker is immune from criticism or social consequences. If I call you a bunch of offensive names here, I'll get downvoted for sure. The comment might be hidden from most. I might get shadowbanned or totally banned, too. That was true of private spaces long before HN existed. If you're a jerk at a party, you might get thrown out. I'm sure that's been true as long as there have been partie…

> The only thing "the SJW crowd" has changed is which words are now seen as offensive.

Well, that, and also bullying thousands of well-meaning projects into doing silly renamings they didn't want or need to spend energy on. Introducing thousands of silly little bugs and problems downstream, wasting thousands of productive hours.

Post reply on HN