Live data from Hacker News

Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

daringfireball.net

771–776 of 776 posts

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#772
post #709

> "The exact words we choose when writing matter." Then write your own damn text if you care about the exact wording so much

You are responding to the original article, authored by John Gruber, who, last I checked, has been writing his own damn text multiple times every damn day for multiple decades as a primary vocation. I’m willing to venture that Gruber is on the list of folks that get to hold the opinion choice of words matters.

Holding that opinion is fine; holding that opinion while using LLMs to probabilistically choose words for you doesn't add up. You want carefully chosen words? Read a book, or write your own. An LLM is not carefully choosing anything, it's providing words based on probabilities.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#773

> "The exact words we choose when writing matter." Then write your own damn text if you care about the exact wording so much

He's a professional writer. As the article says, even if you write your own words, this is still a problem with proofreading, copy-pasting references or quotes, and "AI checkers".

I still don't get why a professional writer would ever accept an AI proofreader over a real one, or a "quote" returned from an LLM over one sourced from a reputable reference.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#774
post #650

> “By definition it must make text worse … because the nature of the watermarking algorithm requires it to sometimes increase the probability of selecting a worse word choice and decrease the probability of selecting the model’s best choice.” Gruber made an effort to but doesn't fully understand how SynthID works. LLMs select the next word randomly from a set probability distribution, so there is no "best choice" unl…

> LLMs select the next word randomly from a set probability distribution, so there is no "best choice" unless you run the LLM with a temperature of 0 which would give out terrible results. I'm not understanding how the word with the highest probability isn't the "best choice"?

Because it's pure exploit on the explore/exploit tradeoff. The best outputs come when the LLM comes up with lots of different ideas, considers them, and selects the best one. If you sample at low temperature it tends to regenerate the same ideas over and over.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#775

Earlier quoted context omitted.

You know, back in the era when proofreaders were human, I never one met a proofreader who rewrote my text afresh, rather than annotating the text with a pen. It's still possible to use Claude to proofread - highlight grammatical, flow, structure, logic errors and make simple suggestions for you to pick and choose or adapt as you wish. No watermarking will flag your text. No flaw accusations of LLM authorship will hau…

Anthropic's own explanation ( https://www.anthropic.com/news/claude-text-watermark ) is not that clear-cut: "When Claude proofreads text written by a person, what it gives back has generally only been lightly edited; because nearly all the words are the person’s, there’s very little (if anything) for the watermark to attach to. Depending on the length of the text and how heavily Claude has edited it, those changes mi…

I think they should release metrics for "how much of this is LLM-written" (the raw test statistic, not including length) along with detection.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#776
post #555

Earlier quoted context omitted.

Interestingly, in a test c't magazine did a long time ago, the test audience preferred mp3 256kbit over the original.

Could be c't magazine accidentally played the mp3 version louder. Human's have a known preference for louder music, and will tend to prefer louder samples over quieter samples. Rumor in the industry is that this was a trick MS used to try to push the WMA format, that they encoded some WMA samples used in some publicized tests at +3dB above the source sample.

They were thorough. It‘s in their 6/2000 magazine.
Post reply on HN