Would this work as an attack? 1. Get a pornographic picture involving young though legal actors and actresses. 2. Encode a nonce into the image. Hash it checking for CSAM collisions. If you've found a collision go on to the next step, if not update the nonce and try again. 3. You now have an image that, to visual inspection will appear plausibly like CSAM, and to automated detection will appear like CSAM. Though, pre…
So at this point we have an image that computers think is CSAM and people think is CSAM, and when held up next to the original verified horrific image everyone agrees is the same image. At this point, someone is going to ask, rightly so, where that came from.
In order to generate this attack, you have had to go out and procure, deliberately, known CSAM. Ignoring that it would be easier just to send that to the target, rather than hiring talent to recreate the pose of a specific piece of child porn (or 30 pieces to trigger the reporting levels), the most likely person by orders of magnitude to be prosecuted in this scenario is the attacker.