Earlier quoted context omitted.
> I would want to hear your explanation for that but more importantly, The people who want it need to prove that it will help. I had started in to list paragraphs and paragraphs of failure modes, unintended consequences, and directly bad effects, but then I realized that that implicitly accepted an inappropriate burden of proof. I am tired of hearing "Something must be done. This is something. Therefore this must be…
> After you've given me that, I'll be happy to rip it to shreds for you. Rip this to shreds: if at least one victim of a crime can be prevented by catching a witless criminal then it is worth it. Also this: Homomorphic encryption and differential privacy are a thing, if certificate transparency logs can prove no certs are issued maliciously then similar logs can be published publicly to prove that the scanning list (…
You didn't come close to what I asked for. You gave no details on how even one crime would be prevented. But I'll give you that it probably would prevent a few.
It's still crazy to say that outweighs any other consideration.
You have to prevent more evil than you cause. That includes the opportunity cost of the resources you put into it, by the way.
"If it saves just one child" is a childish argument and I'm not going to engage with it any further.
... and if you think no children will be sexually blackmailed using the reports you're trying to generate, you are insane.
> Also this: Homomorphic encryption and differential privacy are a thing,
You seem to have slipped from the idea of client-side scanning into using homomorphic encryption to run ML in zero knowledge on the server side. That won't work.
There's a sharp limit on how many operations you can do on the encrypted data before you lose the ability to recover the result. That limit is NOWHERE NEAR the number you need for any useful ML. Plus of course the incredible compute, communication and storage load.
Although I have a general feel for what you can and can't really do in zero knowledge, I am not an actual expert on that technology. So here's a link to an actual expert: https://blog.cryptographyengineering.com/2023/05/11/on-ashto...
I have no IDEA how you'd expect to use differential privacy for any of this. And I very much doubt that you do either.
So let's stick to client-side scanning, since we at least know how to build that. Anyway, the fatal problems are with scanning in general, not with any particular way of going about it.
> if certificate transparency logs can prove no certs are issued maliciously then similar logs can be published publicly to prove that the scanning list (similar to a CRL) does not contain hashes/patterns that haven't gone through appropriate legal/regulatory checks and balances.
I think you may now have moved again, to perceptual hashes. But I guess maybe you could do something like that for an ML model too.
The "checks and balances" I've seen so far have been not so much crappy as nonexistent. Including having private groups create the lists with no real oversight.
Once the infrastructure was in place, I'd expect some countries to PUBLICLY AND OPENLY expand the scope, so auditing is irrelevant anyway. In fact, the OSB already covers more than CSA and would probably require scanning for more than CSA.
> E2EE can also be used to secure the message between the gov agency and the device.
These proposals, especially the OSB, generally call for platforms to manually vet the "hits" before they go to any government agencies, or even to the private advocates.
They kind of have to, since--
1. Most of the government agencies are too swamped to actually follow up on most of the reports they already get, and
2. The bills all demand going beyond looking for specific, already known and vetted files, into looking for things that "look suspicious". Once you go there, the number of false positives will be more than the number of true positives. Especially when you make them terrified to have any false negatives.
But sure, you can encrypt the stuff at each in-motion hop. Which has nothing to do with the main exposures.
> Certificate authorities already have similar exposure that can compromise all your traffic (including the signal app download/install),
That's one reason high-security applications don't trust CAs. Signal doesn't, for instance, but Signal's not special in that way.
> No, the message you have been spreadning is that it is impossible to secure a message and spy in it which you just admitted there.
Well, yes, it is. At least to any reasonable standard.
> So can signal's source code repo infrastructure.
One risk doesn't justify taking on another risk. Especially not a much greater risk.
> It is absolutley possible to implement a scanning infrastructure that has the same security properties as the app's code or app store download/signature security.
OK, this is truly insane.
One system tries to keep messages tightly compartmented, exposing them only to their senders and recipients. It keeps the messages encrypted except on the senders' and recipients' devices, and actually being read.
Another system does the same, except that it also sends some of those messages into a central database. It does so for the purpose of having them read by third parties, and by this I mean humans. In practice, that database has to have a long retention time and a huge number of authorized users. If those users decide that the messages were true positives, they get forward into yet another database.
The second system has every exposure the first system has, plus a bunch of other, worse exposures. It vastly expands the spatial and temporal areas where sensitive data are kept. It puts all users and devices' data into the same compartment. It has probably more than twice the total code of the first system. It trusts thousands more people. And the database is not only huge, but rich in abusable material, so it's a gigantic target that will attract attackers.
Those two systems do not have "the same security properties". They don't have anything close to "the same security properties". Using the phrase "the same security properties" anywhere near those two systems shows that you know nothing about what you're raving about.
And, yes, that second system is what you will get.
> If you are claiming the gov end can be abused by malicious humans then that is beyond your exertise to police humans breaking laws, so long as transparency logs can be produced to criminally punish violators.
Not only am I entitled to my opinions as a member of the (world) polity, but that as an actually competent security specialist, I do have real expertise in designing security systems around how humans actually act.
Unlike you, obviously.
> I am sure you are aware that it is possible to scan for messages without sending off a copy of the message off device and also without informing anyone of false positive hits.
It's a false positive because you don't know it's not a true positive. You therefore have to treat it exactly the same as a true positive.
> Or even contents of true positive hits (requiring them to get a proper warrant and target the device for intrusive collection).
You won't meet a probable cause standard with those hits. I'm sure you could get a warrant in Iran.
Nor do most law enforcement agencies have the resources to get warrants and raid people "on spec" like that.
... but in fact nobody's going to try to build that. The "workflow" you will actually get with this stuff is
1. Device gets a hit.
2. Device sends the data the hit was based on to the platform
3. Optional, but likely to be common because it limits risk for the platform: automation disables the user's account until the platform gets around to reviewing the hit. This may take weeks, especially if some score is borderline.
4. Platform employee reviews the hit (with little context, which matters a lot especially for the text scanning people are demanding)
5. If the hit looks criminal, platform employee forwards it to law enforcement or whoever. With all the data unless legally prevented.
6. If the hit looks borderline, or the user looks like they "might be a risk", or maybe even like they "might generate a bunch more false positives we have to review", platform employee ends the business relationship with the user.
7. If the hit looks completely false, platform employee reenables the user's account.
On hits with very low scores, you might have the device just refuse to send the message or whatever. In that case, either the innocent user is screwed, or the guilty user tries other ways until one works.
> Only active and imminent harm to humans is a reason for civil disobedience not mere speculation and disagreement.
We're seeing people disappeared constantly in a bunch of countries (not the UK so far). I'm OK with calling that active harm.
We're seeing stuff like teenagers jailed for sexting in the US... which this nonsense would definitely greatly increase. And we're people worldwide driven off of platforms. They "look too much like" abusive users to a computer, you see... especially to a computer programmed by some clueless whitebread idiot.
> You are on the side of the oblivious technocrats profiting from harming people.
That stupid bullshit again. News flash: CSA is not profitable for platforms. They don't get paid for it, and if it's visible it drives away profitable users. At MOST it's profit-neutral, usually negative.
> Here I am as your peer engaging with you while being aware of most of the technical facts and you are not convincing me.
That's because you're an obvious fanatic.
> Your strategy of using technical expertise to deceive people
Nobody's tried to deceived anybody. Except maybe you.
> downvoting people like me attempting to engage in civil discourse with you in good faith.
I haven't downvoted you. I have wasted my time engaging with you. I'm not going to waste any more, though.
> I think you will just end up facilitating the building of "GFW" level national firewalls in the west with your techniques.
I'm sure that'll be good for your agenda.