Live data from Hacker News

Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

daringfireball.net

161–170 of 776 posts

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#161
post #68

Earlier quoted context omitted.

This is the classic misunderstanding that LLMs only pick the next token at a time. Really, they are coalescing the probabilities of a range of tokens at a time. There is no “oops, I wrote ‘th’ but I should have written ‘tw’ so I guess I’m stuck writing three instead of tween”.

>There is no “oops, I wrote ‘th’ but I should have written ‘tw’ so I guess I’m stuck writing three instead of tween”. You're mixing up two claims here, and only one of these is kind of true. Yes LLMs do internally plan ahead in a way that is emergent rather than strictly part of their architecture, so that part of your claim is true. The way you word it by saying they are "coalescing the probabilities of a range of t…

What a crazy link:

  So we train a second copy of Claude to work backwards—reconstruct the original activation from the text explanation. We consider an explanation to be good if it leads to an accurate reconstruction. We then train Claude to produce better explanations according to this definition using standard AI training techniques.
Incentives to train a pathological liar. There's no baseline so can only catch out the worst of the lies/errors. Anything (including fabrications) that passes our filters is reinforced?

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#162
post #82
post #74

Earlier quoted context omitted.

It didn’t take, apparently.

It’s fully possible I didn’t explain it very well in the first place, but he is making a wider point. The point I made (quite briefly) is that watermarking is only feasible because for good writing it is necessary to use T>0, or the writing will never explore a more creative choice, and that at T=0 you don’t even need a watermark to spot LLM-generated text. The point he is making is consistent with this, isn’t it? Ei…

The point he is making is not consistent with understanding how temperature influences LLM text generation, no.

He repeatedly states that choosing "the best word" is the most important thing to him. I don't know how you reconcile that with creativity itself, let alone probabilistic sampling.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#163
post #77

Earlier quoted context omitted.

This is not it, no. He is not using AI and it is not I think remotely in his nature to surrender that control. He is engaging with this on principle. Again I am not sure I agree with him, but then it’s a hypothetical because I am not going to get an LLM to write for me either.

Well, it cant be that he is super worried on behalf of people who publish AI slop. That’s not a credible motivation. In fact, he complained a lot about the new ChatGPT app so I can’t believe your claim that he is not using AI. Seems like he really likes to use LLMs and is worried that quality will be degraded. But he will never demonstrate such degradation scientifically, we don’t have anecdotes even.

His complaint about the ChatGPT app is that it’s a shitty non-Mac-ish Mac app. Complaining about shitty non-Mac-ish Mac apps to people who hate shitty non-Mac-ish Mac apps is more or less how he became a full time writer.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#164
> because the nature of the watermarking algorithm requires it to sometimes increase the probability of selecting a worse word choice and decrease the probability of selecting the model’s best choice.

... and the opposite is also true, sometimes it will increase the probability of choosing the "best" word choice. So watermarking makes the LLM quality better then? /s

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#165
post #82

Earlier quoted context omitted.

It’s fully possible I didn’t explain it very well in the first place, but he is making a wider point. The point I made (quite briefly) is that watermarking is only feasible because for good writing it is necessary to use T>0, or the writing will never explore a more creative choice, and that at T=0 you don’t even need a watermark to spot LLM-generated text. The point he is making is consistent with this, isn’t it? Ei…

The point he is making is not consistent with understanding how temperature influences LLM text generation, no. He repeatedly states that choosing "the best word" is the most important thing to him. I don't know how you reconcile that with creativity itself, let alone probabilistic sampling.

I mean creativity there in the LLM sense (temperature driving more creative solutions), not the human sense, and in the context of its consistency, configurability and being amenable to analysis. The watermarking approach makes that non-reproducible, yes? Because Anthropic can and will change it as they see fit.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#166

Respectfully, you are all missing the point. Watermarking is bad not just because of the principled stance that your tool should not be working against your own interests (the passionate argument in TFA), but specifically because it lends credence to the idea that AI detection is a valid and possible thing to do perfectly . As technologists of course we know "oh well yes but with some confidence interval we can detec…

You should submit an article about this instead of having your point buried in a comment section of an article making an unrelated argument.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#168
post #19
post #9

> I want any LLM I use to choose the very best, most precise words at every single decision point. Then bad news: LLMs already use randomness in a fundamental way. Each time they go to generate a token, they first generate a probability distribution of possible tokens. Then they pick one randomly according to this distribution. The technique described can be thought of as making the random number generator pseudo ran…

I think this is a key reason why humans write better prose than LLMs - we can try to choose the best word every time, and go back and restructure sentences and paragraphs if we want. On the other hand, LLMs are forced into picking some likely-ish word, and then have to build the rest of their response to retcon that choice into making sense. Even good human writers would probably struggle with this constraint. It wou…

Humans already do struggle with this constraint. Good examples are JRR Martin, Tolkien, and Rothfuss. You cant describe the struggle of picking the next word and then act like humans don't sit at the table struggling to pick the next word.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#169
post #98

Earlier quoted context omitted.

I'm not sure the quoted statement is true. Proofreading like "point to problems in the text", if you fix the problems yourself and don't copy-paste the solutions given to you, should still be safe, shouldn't it? So, human-written text should not be falsely flagged if you use LLM for proofreading. And if you copy-paste the answers from LLM, I think it's only fair the end result gets flagged. You're not writing it your…

> And if you copy-paste the answers from LLF, I think it's only fair the end result gets flagged. You're not writing it yourself. I wonder if it would even get flagged in that case, because wouldn't the probability distribution of a token when the LLM is suggesting an edit to your writing be different than the distribution of that token once it is in the context of the text it's editing?

I don't even ask LLMs to go that far. Tell me if I've made a spelling, punctuation, or grammatical error, period. Don't rewrite a thing, because LLMs suck at that.
Post reply on HN