Live data from Hacker News

Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

daringfireball.net

691–700 of 776 posts

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#691
post #650

> “By definition it must make text worse … because the nature of the watermarking algorithm requires it to sometimes increase the probability of selecting a worse word choice and decrease the probability of selecting the model’s best choice.” Gruber made an effort to but doesn't fully understand how SynthID works. LLMs select the next word randomly from a set probability distribution, so there is no "best choice" unl…

> LLMs select the next word randomly from a set probability distribution, so there is no "best choice" unless you run the LLM with a temperature of 0 which would give out terrible results. I'm not understanding how the word with the highest probability isn't the "best choice"?

You will understand if you run some LLM models with a greedy sampler that does that. The text quality begins to deteriorate rapidly. This is a very counter intuitive result so I don't blame you for not understanding until you actually tried it and experienced it for yourself.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#692
post #485
post #476

> “He leaped at the chance” and “He jumped at the opportunity” are very similar sentences expressing the same general sentiment, but they are not the same. The exact words we choose when writing matter. I want any LLM I use to choose the very best, most precise words at every single decision point. Neither of these is "better" or "more precise"; in fact, LLMs will generally choose randomly between these candidates ba…

I’m more concerned that this will negatively impact code generation. An additional constraint completely unrelated to code quality is unacceptable as far as I’m concerned. I was an Anthropic user but now I’m looking at OpenAI or even better, open models.

According to their article, these kinds of arbitrary choices don’t come up as often with code, so it’s less likely to have watermarks:

https://www.anthropic.com/news/claude-text-watermark#:~:text...

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#693
I just don't like the idea of an AI company saying that they've added a secret signature that isn't verifiable by any third party to all responses. And we're supposed to both take them at their word, and feed them all the content we want to check so they can continue gobbling up a bunch of fresh works. What is to stop them from saying "Oh yeah, that is ours. We signed it. Trust us."

The obvious conflicts here are wild.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#694
post #165

Earlier quoted context omitted.

The point he is making is not consistent with understanding how temperature influences LLM text generation, no. He repeatedly states that choosing "the best word" is the most important thing to him. I don't know how you reconcile that with creativity itself, let alone probabilistic sampling.

I mean creativity there in the LLM sense (temperature driving more creative solutions), not the human sense, and in the context of its consistency, configurability and being amenable to analysis. The watermarking approach makes that non-reproducible, yes? Because Anthropic can and will change it as they see fit.

Your response seems to be missing the point completely. Gruber thinks "best" writing is produced by choosing the "best" word (highest scoring token) at each step. This is very clear from his writing.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#695

Earlier quoted context omitted.

> LLMs select the next word randomly from a set probability distribution, so there is no "best choice" unless you run the LLM with a temperature of 0 which would give out terrible results. I'm not understanding how the word with the highest probability isn't the "best choice"?

You will understand if you run some LLM models with a greedy sampler that does that. The text quality begins to deteriorate rapidly. This is a very counter intuitive result so I don't blame you for not understanding until you actually tried it and experienced it for yourself.

> You will understand if you run some LLM models with a greedy sampler that does that. The text quality begins to deteriorate rapidly.

Right, I've done this, and this makes sense to me, but I'm not following how that falsifies the top probability word being the best choice in any particular instance.

"Picking only the best word at each decision point results in a worse final result" seems like an imminently reasonable hypothesis.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#696

Earlier quoted context omitted.

I came here to quote the same sentence. Here's another way to look at it: Suppose there actually is a best word choice. The LLM doesn't know what it is but makes a guess. Maybe it's the best one, maybe it isn't. The probability that SynthID changes the best choice to a worse one is equal to the probability that it changes a worse choice to the best one.

I think that depends on the distribution of good choices and bad ones. There may be 10 choices and maybe 8 of them could be appropriate given a context, and 2 are absolutely nonsensical. Or it could be vice versa. And its a spectrum as well.

But the choices are weighted based on those probabilities. This doesn't affect the weightings, only how the final weighted pseudo-random selection is made.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#698

Earlier quoted context omitted.

Pangram doesn't work.

Like their fp rate is a lie? It works well in my limited testing.

It's idiosyncratic to the point of uselessness in mine.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#699
post #575

Earlier quoted context omitted.

Spot on. Also, in 2003, the war on Iraq, still occupied. And the proxy war on Syria (stifling an unwelcome pipeline project). European “leaders” pretend to not comprehend how they're being screwed. Stockholm syndrome. Populations don't understand, propaganda (“free press”) working correctly.

It's very weird looking in from the outside. I mean, sure, the US is the dominant power and everything, so things like Iraq and Iran could be construed as just collateral damage from their imperial maneuvers. But it has just piled on, more and more, and even when they are hit right in the face with the massive bombing of NordStream, still practically nobody tries to draw any sort of line. Reminds me a bit of that 'Ye…

« right in the face » — It's mainly Germany that's hit right in the face and where the Stockholm syndrome applies most closely. If we look at our dear neighbours, then we can see that some of them stand to profit from the Nord Stream bombings. Poland, of course, but also the honorable Norway (have a lot of gas) or the Netherlands (have the most important port). These two also happen to be more closely allied and aligned with the (F)UK/US complex, not directly “Five Eyes” (Anglo only), but “Nine Eyes” (plus France and Denmark). The war had been actively prepared since at least 2014 (military “#TrainToWin” etc), so there was enough time for scheming and dealing.

Guilt with a corollary of subservience to “Western values” having become the predominant ideology in Germany, strategic shortsightedness and failure to properly “relaunch” after 1990, also the occasional murder of more “promising” members of the political spectrum (sans doute at the hands of our dear “friends”), brain drain to the U.S., catastrophic “investments” like Chrysler or U.S. telcom, and barely anything of it ever raised to the threshold of public debate and intelligent reflection : a lot of things combine to explain the dismal state of affairs in modern Germany.

It cannot be ruled out that the German gov was complicit in the bombings; not wilfully, but passively, like someone too weak to resist for lack of self-esteem. Like chancellor Scholz on Feb 7, 2022, at the White House press conference where Biden threatened the pipeline:

▪ “If Russia invades, that means tanks and troops crossing the border of Ukraine, again, then there will be no longer a Nord Stream 2, we will bring an end to it.” ▫ “But how will you do that, exactly, since the project is within Germany’s control?” ▪ “We will… I promise you we will be able to do it.”

There was no reaction from Scholz. Just nothing. By the way, could be I'm wrong, but that part of the press conference seemed scripted/scheduled to me. Not the exact words, but the contents.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#700
post #639

Earlier quoted context omitted.

> So I agree that the watermarking has a cost. But you can't leave out that it is an attempt to reduce negative externalities of AI. Whether it's a realistic or worthwhile attempt is a whole other debate (and Gruber does a good job of debating just that in the latter part of the essay), but saying that the user's needs are the only thing that should ever be considered is reprehensible. Computers are tools that exist…

I agree with you as long as things are at a small scale. But at a large scale, these things reshape society and what it means to be human. Social networks started out[1] as almost wholly good. The negative effects came from scale. Same with advertisement-funded websites and tons of other things that started out as being overall positive for the commons and ended up being highly negative. You can't just stick your fin…

I just don't agree, and that's ok.

I love your point about social media. IMO, initially they didn't know what they had with social media - it took the right kind of uninhibited sociopath to monetize/weaponize it. That's why the MySpace guy is traveling and taking photos while Zuck is building island lairs.

This is where a functional government would step in. There's no magic in social media technology or LLMs. I'd think of them as a car. You can buy a Nissan Leaf of a Porsche 911. One is an ok basic car. The other is an (over) engineered experience in the form of a car. You cannot legally drive a Porsche to it's potential on the public highways... because we have laws that regulate driving and hold the operator accountable.

We're thinking about Claude Code or Gemini or whatever. The people running these companies are like "I want to exceed the power of John D. Rockefeller or Stalin." If we or the EU are going to regulate AI Labs, you need to grab them by the throat and they should be screaming about it. They seem to be very pleased with themselves. Barring real regulatory teeth, we would be opening the aperature to the Chinese companies to rationalize the valuations.

Post reply on HN