Live data from Hacker News

Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

daringfireball.net

81–90 of 776 posts

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#81

Watermarks are garbage because they may embed account id, IP address and deanonimize you. That's why we should be using open-weights LLM whenever possible.

This is the first post I've seen mention it. How traceable are the embedded codes?

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#82
post #74
post #71

Earlier quoted context omitted.

> Does the author think he is currently getting T=0 output from Claude? Is he under the impression that T=0 produces the "best" writing? No and no. I am not sure I agree with his point but I know he is not ill-informed on either of these points, because I mentioned them to him a couple of days ago.

It didn’t take, apparently.

It’s fully possible I didn’t explain it very well in the first place, but he is making a wider point.

The point I made (quite briefly) is that watermarking is only feasible because for good writing it is necessary to use T>0, or the writing will never explore a more creative choice, and that at T=0 you don’t even need a watermark to spot LLM-generated text.

The point he is making is consistent with this, isn’t it? Either you allow temperature to drive creativity, consistently in a way that can be influenced and analysed, or you adulterate that process for the purposes of meeting a corporate/legal directive, in a way that is proprietary and obscure. These are ethically distinct approaches, and since he disagrees with the EU objective he comes down on one side I guess.

Me, I don’t care about the hypothetical enough.

Not least because I think Claude writes depressingly badly and I doubt any steganographic change will enrage me less.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#83
post #12

Earlier quoted context omitted.

Perhaps LLM outputs are uncopyrightable, but derivative works of copyrighted works are not automatically in the public domain.

That's an intriguing twist, isn't it? It could lead to a tug-of-war. Working backwards: if it is possible to confirm 100% confidence that a chunk of text is LLM output, then it is "PD until proven otherwise". How can a human reliably assert human authorship of their source text? When all watermark tests fail? Is that proof of humanity now? If a human proves human authorship, and LLM watermarking tests positive, then…

[deleted]

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#84

This article feels slightly incoherent. You want high quality precise writing and to use an LLM to generate it? Feels like those are diametrically opposed

Exactly. The watermark is proportional to how much text is AI generated. Either the AI really just “fixed some typos” (not enough AI content to hide a watermark) or the AI did most of the writing (enough AI content to hide a watermark).

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#85
post #78

Crazy how a smart person like this fails to understand the gumbel softmax technique. It does not affect writing quality at all, provably. The very fact that there is generally no "best next token" with 100% certainty is precisely why the trick works (you cannot watermark a response to "respond with the To be or not to be soliloquy from the first folio Hamlet", for precisely this reason).

It seems to me like he started out mad and looked to justify it. I'm skeptical that anybody generating LLM text is really all that concerned about optimal word choice. Or even particularly good prose. But let's pretend that person exists. If that person tried, say, an open model and that same model with watermarking applied, I'd be eager to hear their thoughts on the prose quality. Especially if they built an experim…

Google has A/B tested watermarking on millions of responses. They say they observed no difference in user behavior.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#86
post #39
post #19

Earlier quoted context omitted.

I think this is a key reason why humans write better prose than LLMs - we can try to choose the best word every time, and go back and restructure sentences and paragraphs if we want. On the other hand, LLMs are forced into picking some likely-ish word, and then have to build the rest of their response to retcon that choice into making sense. Even good human writers would probably struggle with this constraint. It wou…

This was true in the ChatGPT era. Now we're in a world with reasoning tokens, where a model can thoroughly plan out the response it wants to make. If anything, it makes the style worse.

Yes, models can reason and plan, which helps them write more coherently. But when they write the final output, it’s still a single generation. It would be like letting a human make notes and write an outline, but not let them use the backspace once they start typing their response.

Presumably you could use the same reasoning trace, run multiple generations, and get different outputs (if the temperature is >0).

But now I’m interested in playing more with Cowork or Claude Code/Codex for prose writing to see if the set of tools there affects outputs at all. I guess you might need a more custom “writing” harness.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#87
post #73

Crazy how a smart person like this fails to understand the gumbel softmax technique. It does not affect writing quality at all, provably. The very fact that there is generally no "best next token" with 100% certainty is precisely why the trick works (you cannot watermark a response to "respond with the To be or not to be soliloquy from the first folio Hamlet", for precisely this reason).

If this is true then the probability of the detection tools flagging completely human generated text as AI generated is non-trivial. Let's say I write a completely original piece and the detection tool says there is a 36% probability it was generated with Claude. What then? Now it's up to the person looking at the score to cast a subjective judgement. Maybe to me, anything over 25% is unacceptable. Maybe to someone e…

I don't think that's true. I think it's a binary 0% or near 100% probability of a watermark having been detected; the more changes to the text having been made after the text was output by the LLM and the less leeway the LLM had for probable word choices, the longer the passage necessary to see it.

The "problem" is that seeing the watermark doesn't mean that the person claiming to be the author didn't make extensive changes to the output of the LLM, or that the LLM wasn't simply the final editor of something that the author had put a lot of work into.

> Cognitive surrender.

I don't know what this means. It's just drama. Don't let the LLM write for you and this is not a worry. I'm not worried about the poetry of LLM output being subtly adulterated.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#88
I was initially surprised that Gruber was so invested in the "quality" of AI-generated text, which in my mind is an oxymoron. But really, Gruber's interest here is with the EU. This forms part of his ongoing attacks on the EU, all because they have been forcing Apple to align with regulations.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#89
post #73

Crazy how a smart person like this fails to understand the gumbel softmax technique. It does not affect writing quality at all, provably. The very fact that there is generally no "best next token" with 100% certainty is precisely why the trick works (you cannot watermark a response to "respond with the To be or not to be soliloquy from the first folio Hamlet", for precisely this reason).

If this is true then the probability of the detection tools flagging completely human generated text as AI generated is non-trivial. Let's say I write a completely original piece and the detection tool says there is a 36% probability it was generated with Claude. What then? Now it's up to the person looking at the score to cast a subjective judgement. Maybe to me, anything over 25% is unacceptable. Maybe to someone e…

> the probability of the detection tools flagging completely human generated text as AI generated is non-trivial

How does that follow? AI-generated text is already not a perfect emulation of human writing. There's lots of room to affect it laterally without changing the level of quality.

As I understand it, LLMs with temperature >0 can select from many possible outputs. All they're doing is limiting the possible outputs to ones that contain this pattern. I don't see any reason why the quality of that subset should be lower than average. The very best outputs will likely be eliminated, but so will the very worst.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#90
post #36
post #19

Earlier quoted context omitted.

I think this is a key reason why humans write better prose than LLMs - we can try to choose the best word every time, and go back and restructure sentences and paragraphs if we want. On the other hand, LLMs are forced into picking some likely-ish word, and then have to build the rest of their response to retcon that choice into making sense. Even good human writers would probably struggle with this constraint. It wou…

autoregressive generation doesn’t mean the model is myopic. the next-token distribution can already reflect a longer horizon plan for the output sequence.

Sure, but mightn’t there be several plausible long horizon plans?

Here’s an example: I had asked Claude for some music recommendations in a certain style. Part of its output was:

*Long journey tracks*

Clinic — “The Return of Evil Bill”

Guided by Voices — not really, wrong band

Silver Apples — “Oscillations”. Proto-everything, deeply repetitive, hypnotic.

So at some point there, the next token produced was “Guided” or “Guide” or whatever, and then because it can’t go back, it had to correct itself after the fact.

Reasoning/CoT have helped a lot, but I feel like small versions of this still happen all the time.

Human writing is like 90% editing.

Post reply on HN