Live data from Hacker News

Show HN: Neural-hash-collider – Find target hash collisions for NeuralHash

github.com

271–280 of 363 posts

Re: Show HN: Neural-hash-collider – Find target hash collisions for NeuralHash

#271
post #254

Earlier quoted context omitted.

Images, leaked government files, anti-administration phrases, unflattering memes of the president, statements that contradict the government's current stance, etc. All things current at social media companies seem willing to censor after suggestion of the administration.

You seriously think Apple would voluntarily search private devices for images which aren't illegal and don't even hint at any action which is illegal? I don't think you're being serious.

Why not? Facebook and Twitter have done exactly that in the past year. Why is it far fetched for Apple suddenly, given this amazing reverse-course on branding?

The only realistic alternative to Apple is Android... And Google is pretty darn transparent in their spying on users. Apple just did a 180 degree about-face on all the branding they've built over the last decade. Why should anyone trust Apple again?

Look, this whole neural-hash thing took what, 2 weeks for people to fabricate collisions? This just illustrated how poorly conceived and ill-thought the entire plan was from Apple. It's not beyond reason to assume any of these things given the evidence we currently have.

Re: Show HN: Neural-hash-collider – Find target hash collisions for NeuralHash

#272
post #100
post #82

Really naive question. What's to stop apple from using two distinct and separate visual hashing algorithms? Wouldn't the collision likelihood decrease drastically in that scenario? Again, really naive but it seems like if you have two distinct multi-dimensional hashes it would be much harder to solve the gradient descent problem.

I'm fairly sure they do, actually. It was in one of the articles earlier today that Apple has a distinct, secret algorithm they perform on suspected CSAM server side after it gets flagged by the client side neural hash. Then only after 30 such images from a single user are identified as CSAM by both algorithms will they be sent to a human reviewer who will confirm their contents. Then, finally, law enforcement will b…

But relying on an algorithm staying secret is security-by-obscurity 101. You can rely on a cryptographic key staying secret; you can't rely on the design of an algorithm staying secret (I do agree there's a little blurring of these lines with large, trained models, but the gist remains - you can't just hope that nobody sees the structure/weights of your second hash function).

Re: Show HN: Neural-hash-collider – Find target hash collisions for NeuralHash

#273

Earlier quoted context omitted.

idk what's in the database, whether it's rape or nudes or both. Although depictions of sexual acts versus simple nudity seems like a logical place to draw a line, all the lines on adult pornography are arbitrarily drawn based on "community standards", and we're only a few decades away from state-level bans on any nudes as "porn" in the US, including artistic photos. (Not to mention anti-sodomy laws). Even if what's i…

You would presumably have the 30+ images on your device or in iCloud to prove your innocence. For you to get caught up in this dragnet, 30+ plus images have to match NeuralHash’s of known illegal images, thumbnails of those images have to also produce a hit when run through a private hash function that Apple only has, and two levels of reviewers have to confirm the match as well.

[deleted]

Re: Show HN: Neural-hash-collider – Find target hash collisions for NeuralHash

#274
post #100

Earlier quoted context omitted.

I'm fairly sure they do, actually. It was in one of the articles earlier today that Apple has a distinct, secret algorithm they perform on suspected CSAM server side after it gets flagged by the client side neural hash. Then only after 30 such images from a single user are identified as CSAM by both algorithms will they be sent to a human reviewer who will confirm their contents. Then, finally, law enforcement will b…

> It was in one of the articles earlier today that Apple has a distinct, secret algorithm they perform on suspected CSAM server side But then they still need to upload the original image to the server, and what was the reason for doing the scanning client-side then when they still upload it?

>and what was the reason for doing the scanning client-side then when they still upload it?

Probably so that China cannot force Apple to hand over arbitrary images in iCloud (or all images in iCloud). With Apple's design the only images China can get from you are malicious images that people send you. If Apple scanned every image serverside without any clientside scanning, then theoretically China could get all newly-uploaded iCloud images.

Re: Show HN: Neural-hash-collider – Find target hash collisions for NeuralHash

#275
post #270

Earlier quoted context omitted.

If a photo is about to be uploaded to iCloud Photo Library then it is scanned for CSAM. If it's not, it isn't. Still waiting on that citation.

Are we arguing the same thing? How does one opt-out a specific photo? It's not possible as far as I know.

I've no idea what your point is. I've tried offering answers for all these random questions, but I'm still waiting for you to offer a citation for the claim you made earlier.

Re: Show HN: Neural-hash-collider – Find target hash collisions for NeuralHash

#276
post #271

Earlier quoted context omitted.

You seriously think Apple would voluntarily search private devices for images which aren't illegal and don't even hint at any action which is illegal? I don't think you're being serious.

Why not? Facebook and Twitter have done exactly that in the past year. Why is it far fetched for Apple suddenly, given this amazing reverse-course on branding? The only realistic alternative to Apple is Android... And Google is pretty darn transparent in their spying on users. Apple just did a 180 degree about-face on all the branding they've built over the last decade. Why should anyone trust Apple again? Look, this…

> Facebook and Twitter have done exactly that in the past year.

You think Facebook and Twitter have dobbed users into the US Government for spreading unflattering memes? You are delusional, or more likely, not being serious. This is the last reply you'll be seeing from me. I'm collapsing this thread and won't be replying any more.

Re: Show HN: Neural-hash-collider – Find target hash collisions for NeuralHash

#277
post #258

Earlier quoted context omitted.

My guess is that internally, they've realized they have a big CP problem on iCloud. That's a huge liability.

I have series doubts about that. CSAM really isn't an issue in the US, culturally and legally.

Really? It isn't an issue?

Re: Show HN: Neural-hash-collider – Find target hash collisions for NeuralHash

#278
post #272
post #100

Earlier quoted context omitted.

I'm fairly sure they do, actually. It was in one of the articles earlier today that Apple has a distinct, secret algorithm they perform on suspected CSAM server side after it gets flagged by the client side neural hash. Then only after 30 such images from a single user are identified as CSAM by both algorithms will they be sent to a human reviewer who will confirm their contents. Then, finally, law enforcement will b…

But relying on an algorithm staying secret is security-by-obscurity 101. You can rely on a cryptographic key staying secret; you can't rely on the design of an algorithm staying secret (I do agree there's a little blurring of these lines with large, trained models, but the gist remains - you can't just hope that nobody sees the structure/weights of your second hash function).

Security by obscurity has protected a lot of things successfully...

Re: Show HN: Neural-hash-collider – Find target hash collisions for NeuralHash

#279

Earlier quoted context omitted.

Sorry, this is not even wrong. The visual derivative is just a resized, very-low-resolution version of the uploaded image. "Matching the visual derivative" is completely meaningless. The visual derivative is not matched against anything, and there is no "original" visual derivative to match against. If enough signatures match, Apple employees can decrypt the visual derivatives, and see if these extremely low resoluti…

I just want to be clear if I understand this... many images can result in the same hash, but the hash can and will be reversible into one image? And that image is a low res porn photo derived from the algorithm's guesswork? So once a hash matches they don't check if there was a collision and the photo is completely unrelated, they just see the CG porn? If that's the case then why even look at the derived image?

No, this is not what's going on at all. The employees never see the original photos in the government CSAM hash database. Apple doesn't even have these photos: it's precisely the kind of content that they don't want to store on their servers. If some conditions are satisfied, the employees gain access to the visual derivatives (low-resolution copies) of your photos, and they judge whether these look like they could plausibly be related to CSAM materials.

The exact details of the algorithm are not public, but based on the technical summary that Apple provided, it almost certainly goes something like this.

Your device generates a secret number X. This secret is split into multiple fragments using a sharing scheme. Your device uses this secret number every time you upload a photo to iCloud, as follows:

1. Your device hashes the photo using a (many-to-one, hence irreversible) perceptual hash.

2. Your device also generates a fixed-size low resolution version of your image (the "visual derivative"). The visual derivative is encrypted using the secret X.

3. Your device encrypts some of your personally identifying information (device ids, Apple account, phone number, etc.) using X.

4. The hash, the encrypted visual derivative, and the encrypted personally identifying information are combined into what Apple calls the "safety voucher". A fragment of your key is attached to the safety voucher, and the voucher is sent to Apple over the internet. The safety vouchers are sent in a "blinded" way (with another encryption key derived using a Private Set Intersection scheme detailed in the technical summary), so that Apple cannot link them to specific files, devices or user accounts unless there's a match.

5. Apple receives the safety voucher. If the hash in the received safety voucher matches that of known CSAM content in the government-provided hash database (as determined by the private set intersection scheme), the voucher is saved and stored by Apple, and the fragment of your secret key X is revealed and saved. (You'd assume that they filter out / discard your voucher if there's no match; but the technical summary doesn't explicitly confirm this; this means that they may store and use it in the future to run further scans).

6. If your account uploads a large number of matching vouchers, then Apple will gather enough fragments to reassemble your entire secret key X. Now that they know your secret key, they can use it to decrypt the "visual derivatives" stored in all your saved vouchers.

7. An Apple employee will then inspect the "visual derivatives", and if your photos look like CSAM (more precisely, this employee can't rule out by visual inspection that your photos are CSAM-related), they will proceed to use your secret key X (which they now know) to decrypt the personally revealing information contained in your safety voucher, and report you to the authorities.

Keep in mind that the employee looking at the visual derivative does not, and cannot, know what the original image is supposed to look like. The only judgment they get to make is whether the low-resolution visual derivative of your photo looks like it can plausibly be CSAM-related or not. Plainly speaking, they will check if a small, say 48x48 pixel, thumbnail of your photo looks vaguely like naked people or not.

Re: Show HN: Neural-hash-collider – Find target hash collisions for NeuralHash

#280
post #50
post #49

Earlier quoted context omitted.

The FBI off their back that they aren’t doing enough to stop the spread of CP.

I find it hard to believe CSAM was so pervasive on iDevices that they'd feel compelled to do something about it. As far as we know (and I'm sure lots of eyeballs are looking now) Android doesn't do this. And frankly, why would Apple care that the FBI isn't cozy with them. Their entire brand is "security and privacy", kind of goes against most 3 Letter Agencies anyway.

At least according to [1]:

"Last year, for instance, Apple reported 265 cases to the National Center for Missing & Exploited Children, while Facebook reported 20.3 million, according to the center’s statistics. That enormous gap is due in part to Apple’s decision not to scan for such material, citing the privacy of its users."

If you were a law enforcement agency and noticed this discrepancy, would you believe that you'd be letting some number of child abusers get away because of that difference in 20 million reports? iCloud probably doesn't have the same level of adoption as Facebook, but the gap is still very large.

[1] https://www.nytimes.com/2021/08/05/technology/apple-iphones-...

Post reply on HN