I too wonder whether cryptographic signatures will be the long run solution to deepfakes. Can you outline why you don't think that will be the case?
Here's an alt solution to argue against:
1. The necessary PKI gets bootstrapped by social media companies, where deepfakes begin to seriously threaten their "all-in-on-video" strategy, and simultaneously look like an opportunity to layer on some extra blue-check-style verification.
2. Example: you upload your first video to twitter, it sees the AV streams are unsigned, generates a keypair for you, does the signing, and adds the pubkey half to your twitter account. (All of this could be done with no user input.)
3. The AV streams have signatures on short sequences of frames. Say every 24th video frame has a signature over the previous 24 frames embedded into the picture. Similarly, every second of audio has a signature baked in.
4. The signatures aren't metadata that can be easily thrown away. They're watermarked into the picture and audio themselves in a way that's ~invisible & ~inaudible, but also robust to downsampling & compression. This is already technically possible.
5. Since we're signing every past second of the AV streams, this works for live video.
6. Viewers on the platform see a green check on videos with valid signatures; maybe they even see the creator's twitter handle if this is a reshare.
7. Like all social media innovations, the other major platforms copy it within the next 6 - 12 months. POTUS uses it. People come to expect it.
8. Long run: the public comes to regard any unverified video footage with suspicion.
Why won't this deepfake solution come to pass in the long run?