Live data from Hacker News

Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

daringfireball.net

501–510 of 776 posts

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#501

I’m surprised by the comments here being so favorable to anthropic. The comments are right about there being no “best” token, and yeah Gruber may have an agenda here. But I think the fundamental principle is that this approach messes with the distribution in ways that deviate from the trained model. Take the “gray” and “overcast” choices. And lets say before applying synthid the percentages were 48% and 52%. Those pe…

> But I think the fundamental principle is that this approach messes with the distribution in ways that deviate from the trained model. This. I want the model I'm paying for to be "pure". I don't Anthropic or anyone else messing around with it, especially not for idiotic reasons like facillitating AI stigmatization. The "safety" nonsense is obnoxious enough. They should train the best possible model and let the weigh…

> I want the model I'm paying for to be "pure".

I don't think the models are pure in any meaningful sense. The labs have some idea of what kind of output they want from the models and then they put a huge amount of effort into training the models on the right sorts of data and massaging the models afterwards to push them towards the desired output. Then at a more practical level there's the layers of filters before your prompt even hits the model (e.g. anthropic's auto-mode classifier), system prompts, response level filtering etc.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#502

Earlier quoted context omitted.

> But unless you know the prompt, you don't fully know which choice the model faces. Surely a lot of coding space is wasted compensating for that uncertainty. When you only need to encode one bit, the signal to noise ratio can be very low. If I try to write my own human words under the policy of "try a little bit to avoid the letter 'e' in every fifth word," then a sufficiently long text would still be 'watermarked'…

Whereas if you fully avoid the letter e, everyone will know you are George Perec

Who is Gorg Prc?

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#504

I’m surprised by the comments here being so favorable to anthropic. The comments are right about there being no “best” token, and yeah Gruber may have an agenda here. But I think the fundamental principle is that this approach messes with the distribution in ways that deviate from the trained model. Take the “gray” and “overcast” choices. And lets say before applying synthid the percentages were 48% and 52%. Those pe…

> I think the fundamental principle is that this approach messes with the distribution in ways that deviate from the trained model.

But it doesn't! The distribution doesn't change at all. The only thing that changes is that sampling of that distribution becomes deterministic as per a precomputed seed.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#505
post #376

Earlier quoted context omitted.

The exact words matter to people who people who don’t use it to write for them. A lot of people use Claude as a friend/therapist/romantic partner/etc. That’s who I imagine would be most affected

>A lot of people use Claude as a friend/therapist/romantic partner People developing a para-social (pseudo-social?) relationship with a corporate robot have far bigger problems than the word-chooser in their robot "friend".

Very intelligent people need intelligent-others to bounce ideas off of, and the LLM can be that.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#506

Earlier quoted context omitted.

That's inaccurate in two ways: (1) The behavior that is approximately what you describe is not "fundamental" (though it may not be something you can disable on some hosted providers), it is an option that is not fundamental (and with runtimes where you have full control can be either disabled or tuned in a large number of manners), and (2) The actual behavior that is approximately what you describe already usually in…

You're generating a pseudo random number one way instead of another way. How would that inherently compromise quality?

1. As watermarked text is added to the training data, watermark-related tokens will be associated more with AI outputs and thus lower quality outputs which will hasten model collapse. Especially because every provider has its own secret key and they are all training on eachother's outputs anyway.

I guess they can at scale filter the watermarked documents (by necessarily allowing eachother to at scale checked for watermarks, but banning the labs not part of the watermarking-cabal). Makes me wonder how useful the human quality filter is on AI output - if a human judges a given output as genuinely good and posts it somewhere for the scrapers to find and take into the training sets, will these types of outputs also be filtered out?

2. (raw, pre-watermarked) Output token probability situations where 1 output token has the majority of the probability mass associated with it, but it is not in the watermarked set, will force the model with much higher probability to walk a non-optimal latent space. E.g., if the next OBVIOUS token for a given sentence would be a point, but the model is in this way not allowed to output it, it might put a comma and start off on a whole different tangent just to make the initial non-optimal comma grammatically make sense.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#507
post #328

Earlier quoted context omitted.

The question here is not to what product management advice to the Gemini team. The discussion here is whether the watermarking is noticeable. The easy thing to do here would be to have 1000 questions, randomly assigning one half to an LLM with a watermark, and the other half without. Then show people pairs and say, "Which one seems watermarked?" (Or, "Which text seems more natural" or "Which is a better answer" or so…

They do exactly that "which is the better answer?" test -- I've seen it pop up a few times.

Isn't "which one is watermarked?" a different question than "which one is better?"

"Which diamonds are shinier, the blood diamond sourced ones or the ethically sourced ones?" ... that's not the same question as "which diamonds are blood diamonds" (to employ an extreme analogy)

Concluding that no one could detect which ones were blood diamonds because they were "equally shiny" is not really correct now, is it?

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#508

Earlier quoted context omitted.

That's an intriguing twist, isn't it? It could lead to a tug-of-war. Working backwards: if it is possible to confirm 100% confidence that a chunk of text is LLM output, then it is "PD until proven otherwise". How can a human reliably assert human authorship of their source text? When all watermark tests fail? Is that proof of humanity now? If a human proves human authorship, and LLM watermarking tests positive, then…

How do you assert it now? I post some text on the internet, you claim you have copyright, how do you prove that?

Copyright law was updated in a very helpful way in the last twenty years sometime so that as soon as you post something to the internet you have copyright. If you need a citation don't hesitate to ask someone else.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#509
post #322

Earlier quoted context omitted.

Interesting, are you being literal about requesting 20 ways of saying the same thing? It seems pretty excessive and I'm actually impressed that asking an LLM to rewrite a statement 20 times yields output that isn't excessively redundant. Does the LLM do a pretty good job reading your mind, or do you still find yourself manually piecing together pieces from the 20 suggestions into a satisfactory sentence?

Yep, I literally ask it to say it in 20 different ways. Sometimes a sentence structure or a combination of words will just work better. The goal is usually to simplify a sentence without losing meaning. I don't expect the LLM to read my mind. The unit of work is too small for intent to matter, and I'll just steer the next recommendations in a direction as needed. Most of the suggestions are crap, but they can contain…

And this beats just writing the piece yourself?

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#510

Earlier quoted context omitted.

EU regulations are going to force it, friend. The tsunami is coming and cannot be stopped and Anthropic has jack to do with it.

I wonder if that is entirely true. It is also in the AI companies' own best interest to be able to detect AI generated text so that they can avoid training on it ("Habsburg AI"). In the case of Anthropic it would also be entirely unsurprising if they've been lobbying the government to force everyone to do something in their (Anthropic's) own best interest.

> It is also in the AI companies' own best interest to be able to detect AI generated text so that they can avoid training on it

It’s also in their interest to demonstrate that they can be trusted and to show that they at least pay lip service to limit the obvious downsides of the tools they are selling. The use cases they sell to mainstream audiences are not affected by detection tools. The point of having a LLM do the work for you is that the work is done, and reliably. It does not matter if it is done by a LLM, and most of the time it is obvious anyway.

Post reply on HN