This also, just another precedent of anti-user, pro Authoritarian, from LLM companies.
Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing
391–400 of 776 posts
Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing
#392Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing
#393Earlier quoted context omitted.
Why not switch it around? Solves problems of privacy (no data upload), is far more doable and it's far more important to be able to verify content has not been tampered with after originating from a human or capture device (like a camera). We also have decades of cryptographic experience in reliably signing text, images, video, etc. and it doesn't break apart because a text is too short. Having proof that content (es…
Thank you for stating clearly situation. I fully agree with your assessment. For almost 10 years now I have been saying that we need to virtually watermark reality. By "virtual" I mean store the metadata about the digital capture on a public blockchain. Then my devices could have a built-in "fake vs real" detector. Artists, photographers, journalist, etc. are going to want and need this.
But even then, people will be able to point that camera at a manipulated/generated image (either printed or on a screen). Maybe that one could be solved if the photo included some depth information?
Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing
#394> I want any LLM I use to choose the very best, most precise words at every single decision point. Then bad news: LLMs already use randomness in a fundamental way. Each time they go to generate a token, they first generate a probability distribution of possible tokens. Then they pick one randomly according to this distribution. The technique described can be thought of as making the random number generator pseudo ran…
Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing
#395Earlier quoted context omitted.
> blindly trusting they won't train on any of that being allowed to train on any data that you can legally obtain ought to be a right for anyone. After all, i am allowed to learn off anything i can legally read (and perhaps even illegally read). The only thing not allowed (rightly so) is to produce a copy with enough similarities that it can be replacing the original.
> being allowed to train on any data that you can legally obtain ought to be a right for anyone. I have the opposit viewpoint to the extreme. They shouldn't be allowed to even read that data until they are very clear about what they will or not do with it. Can they publish it? Can they store it? Can they use the information in it on prediction markets? Etc. Humans reading texts historically come with little negative…
IDK, we do have laws against opening other people's mail. Those have been on the books for hundreds of years. Seems like someone figured out a while ago that certain unauthorized humans reading certain restricted text wouldn't be good.
Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing
#396Earlier quoted context omitted.
>No one uses a pure random function over the whole probability distribution described by the LLM's output. So what? By definition with this system the LLM will chose tokens it otherwise would not, purely for watermarking reasons. Yes this token may have had a decent likelihood of being chosen anyway, but it wouldn't have been chosen and now it was for reasons nothing to do with output quality. I'm not sure what your…
My main point is that sampling with a modified distribution compared to the one produced by the model is already being done, and it is generally found to increase quality, not decrease it. So there is no reason a priori to assume that the watermarked distribution would be lower quality than other schemes for altering the "raw" output distribution (such as top P, top K, temperature, etc). My second point is that the t…
>My second point is that the training of a model by definition maximizes the fitness between the final output function and the training metrics.
Right, but the fitness in question is watermarked text fitness, not fitness for any user interests aligned metric. You're basically saying that if we train LLMs on watermarked text they'll be really good at producing text that looks watermarked, and then we'll stick an actual watermark on top of that. Screw whatever the user wanted it to be good at.
Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing
#397Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing
#398Earlier quoted context omitted.
Why not switch it around? Solves problems of privacy (no data upload), is far more doable and it's far more important to be able to verify content has not been tampered with after originating from a human or capture device (like a camera). We also have decades of cryptographic experience in reliably signing text, images, video, etc. and it doesn't break apart because a text is too short. Having proof that content (es…
Thank you for stating clearly situation. I fully agree with your assessment. For almost 10 years now I have been saying that we need to virtually watermark reality. By "virtual" I mean store the metadata about the digital capture on a public blockchain. Then my devices could have a built-in "fake vs real" detector. Artists, photographers, journalist, etc. are going to want and need this.
Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing
#399Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing
#400Text watermarking is another EU rule made without real world input. The Union is stuck on major economic crises (electricity prices for instance) because nobody can agree on anything. However, the bureaucracy forces tech into a privacy nightmare. Brussels cannot bring together its own members but it loves pretending it can govern the internet.
Except for the input of the hundreds of stakeholders they consulted, Anthropic included [1]?
> The Union is stuck on major economic crises (electricity prices for instance) because nobody can agree on anything.
That sure seems relevant for the implementation of AI watermarking...
> However, the bureaucracy forces tech into a privacy nightmare.
No, this transparency allows consumers to more easily detect AI generated content.
[1] https://digital-strategy.ec.europa.eu/en/policies/code-pract...