Live data from Hacker News

Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

daringfireball.net

331–340 of 776 posts

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#331
post #237

Earlier quoted context omitted.

That cannot be true. The quality of an LLM's output is the quality of the probability calculations for the next token. Anything that degrades the relationship between the system's best assessment of the appropriate probability and the actual probability used is a degradation of the quality of that probability and therefore of the output. If this didn't have a detectable effect on the quality of the token probability…

An encrypted hard drive is EXACTLY uniformly distributed random bytes if you do not know the encryption key. No one would be able to tell the difference between a drive that is just purely random numbers or is actually filled with content. (Of course excluding the usually intentionally added readable header) If this was not the case, the encryption would be broken, and most everyone agrees that good encryption does e…

>You just need to know your LLM distribution and the encryption key, then with each new token you exponentially increase the chance of knowing whether it fits your encryption key. Without actually affecting the token choice in a perceivable manner.

I struggle to understand the relevance of that comment.

The blue/green token list biasing process literally does cause different tokens to be occasionally chosen. Not only that, but because a different token was chosen at one point, this changes the probabilities of all subsequent tokens, and the resulting later token stream every time it happens. If you had access to the token stream as it would have been, and the watermarked one by the end of the text they will be very noticeably different.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#332
What a bunch of entitled whining. How is the system to know that it's just a private conversation that won't be used in some fraudulent way? Abuse is currently rampant, yes please let's find a way to mark LLM output. The thing I'm worried about is giving the providers the power to claim provenance. Even ignoring the privacy issues, the operational hassle of having to check N providers makes these approaches at best limited. I want to see research into providing a shared public or ideally self-hostable oracle that uses some standardized method for watermark detection. Similar to asymmetric crypto where users can't reasonably find out the secret part but can do something useful with it nonetheless.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#333

Earlier quoted context omitted.

I think there are two separate requirements? One that if you post something like an AI video on the internet or anywhere else, you must label it as AI. And another one that AI providers must watermark their outputs. If you get caught uploading watermarked media without the clear label, you're in big trouble, mister.

We were talking about text.

Which must also be marked as AI, I hope. (But I doubt that is the law, because the EU is cucked to big businesses interests)

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#334

Earlier quoted context omitted.

Can we install random unapproved apps on our iPhones yet, or is Apple aiming to just be fined a trillion dollars because they make more than that from the 30% cut?

Yes: https://support.apple.com/en-mk/117767

This says we can only install Apple-approved apps.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#335

Earlier quoted context omitted.

There's no difference between those two things. The distribution that matters is the distribution of tokens that are picked not the distribution of tokens the LLM model passed to the selector.

On average the distribution of selected watermarked tokens is the same as the original distribution. You can test this experimentally. I think there is a possible weakness in the context of the watermarker but that is not your claim iiuc.

The distribution of token "ple" being the same on average, but lower after "crum" and higher after "cou", is not no difference. It's irrelevant that the single-token distribution is unchanged if the joint distribution is different.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#336
post #77

Earlier quoted context omitted.

This is not it, no. He is not using AI and it is not I think remotely in his nature to surrender that control. He is engaging with this on principle. Again I am not sure I agree with him, but then it’s a hypothetical because I am not going to get an LLM to write for me either.

> He is not using AI That's ... even worse? So we're all here in the comments trying to figure out what the author means, and what their overall point is, while clearly they don't even use the damn thing? Oof... What a waste of time for everyone involved.

Why? I really don't understand this. Why can't a tech writer take a deep but neutral intellectual interest in something? Isn't it important that some do? Do people have to be stakeholders or clearly on one given team, pro- or anti-, for their opinion to matter? Is it that tribal?

It seems fully logical to me that someone who writes for a living (who, as it happens, developed the very markup language LLMs use for everything) should be invested in understanding the automatic plagiarism and word calculating machine from an intellectually honest position.

I personally am pretty severely big-two-AI-firms, increasingly anti-big-tech, but I am learning and researching uses of LLMs because for myself I really need to understand how to use them in an intellectually and (as far as is possible) ethically sound way. Learning because as a boring old freelance programmer I have to; foolish to pretend otherwise.

So I completely understand his position — that the AI industry is hot air and crooked and scammy and weird, and some of the people involved genuinely rather dark-sided, but the technology exists and if it hints at threatening your livelihood, you need to understand it.

From reading his work for the best part of twenty years or so (and emailing him intermittently over that time) it would seem to me that he's a lot less bearish on the tech industry than me, and a lot less fond of the EU than I am; he's more optimistic than I am. But he writes because he has to write. I should think that would make him highly invested in understanding what LLMs do.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#337

Gruber was a good voice in the industry but this article misses the mark in a lot of ways. A company the size of Anthropic would not voluntarily jeopardize their massive valuation if they didn’t feel the resulting output would maintain a similar level of quality as before. Is there a similar worry that their system prompt, which is injected at the start of every conversation also influences token generation in an art…

Gruber is not much of a details man - he helped invent Markdown (to be lauded) but ghosted its standardisation. I would be fascinated to hear Prod John MacFarlane of UC Berkeley's opinion on it all given he was heavily involved in the push to get Markdown standardised.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#338
post #246

My biggest concern is that checking any text for watermarks requires sending the entire text to Anthropic. And even that is not sufficient, as the text might have been generated with ChatGPT, Gemini, Grok, Mistral, ... So every check requires sending the text to as many AI providers as offer a watermarking detection API, almost all of which have a very dubious track history with obtaining training data through illici…

Would really like to know how their watermarking technique works, and if they can use it to store arbitrary information in the text (I assume they can if only in a limited way). I assume if the text is long enough they could add all kinds of metadata that would then be undetectable as the model that generated the text is the "cryptographic" key that encodes the data. I wonder if you can have another model run over the text and destroy the watermark. I predict an interesting cat and mouse game to develop.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#339
post #305
post #274

Earlier quoted context omitted.

> blindly trusting they won't train on any of that being allowed to train on any data that you can legally obtain ought to be a right for anyone. After all, i am allowed to learn off anything i can legally read (and perhaps even illegally read). The only thing not allowed (rightly so) is to produce a copy with enough similarities that it can be replacing the original.

> being allowed to train on any data that you can legally obtain ought to be a right for anyone. I have the opposit viewpoint to the extreme. They shouldn't be allowed to even read that data until they are very clear about what they will or not do with it. Can they publish it? Can they store it? Can they use the information in it on prediction markets? Etc. Humans reading texts historically come with little negative…

> Humans reading texts historically come with little negative consequences, but machines reading and processing texts en masse is more dangerous and should be regulated.

Citation needed. This is sounding tautological.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#340

Earlier quoted context omitted.

Seems unlikely unless Anthropic invented a time machine, given that the phenomenon predated Claude 1 by two years, and their stated introduction of watermarking (August 2) by 5.

I see how that was confusingly written. I'm not suggesting Anthropic are somehow retroactively causing it, just wondering if it could create a similar effect.

Less unlikely, but I would still suggest that it is somewhat unlikely for any recent LLM to place any significant probably on, e.g. "disappointment" instead of "failure" as the appropriate token (or series thereof) after "kidney" or any similar example.
Post reply on HN