Live data from Hacker News

Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

daringfireball.net

751–760 of 776 posts

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#751

Earlier quoted context omitted.

The EU mandate for AI watermarking is likely the first step in the direction of prohibiting AI for specific use cases. The pretext for outlawing (or at the very least controlling) the use of AI when the time comes will be something along the lines of data integrity or just general compliance legalese. So, the use cases that are being sold to mainstream audiences actually will be affected by detection tools, especiall…

I’m sorry, this is going to be a bit long but you made good points. > The EU mandate for AI watermarking is likely the first step in the direction of prohibiting AI for specific use cases. I am not sure how practical that would be. The cat’s already out of the bag and they won’t prevent companies in the whole world from releasing open weight models. Playing catch up by distilling flagship models is also relatively ch…

Great points all around, lots to think about. It's totally possible that I am way off in my assessment of the future.

I think your comment about torrenting touches on an interestingly relevant case study, specifically media piracy. The seed of that technology was sown when the internet was still just clusters of machines passing files around and then exploded with the PC, Napster, and TPB. Napster tried to be a legitimate commercial enterprise with a business model of undermining the ability of copyright holders to rent-seek on the consumption of the material they 'owned'. At the core it was a novel technology (P2P) that revealed an economic arrangement to be out of date (if a distributor no longer has to manufacture a copy of the media for each individual consumer, their business model boils down to rent-seeking). Western legal systems were quick to rule on the matter (in favor of copyright holders), and I have no doubt that they were looking ahead to a future wherein media creation was totally disincentivized by said novel technology.

Now a novel technology (the cloud inference-backed LLM) is challenging another economic arrangement. This time around, the arrangement being challenged is the higher education->professional job pipeline. All advertising and media messaging aside, it really does seem like the frontier labs are only economically viable if they get massive enterprise deals across a broad spectrum of industry. There is fundamentally one chunk of capital organizations are going to spend either supporting their talent pipeline or padding it (to put it gently) with enterprise LLM deals. If the latter path is taken too far, consumer spending plummets (due to lack of middle-class incomes), assets backed by consumer debt/spending fail longterm, and we will have to deal with a deluge of socio-political issues stemming from the absence of real social mobility (we are in the early stages of this now, incidentally).

All this to say, I think parallel situations from the past can guide our thinking re: AI regulation and the forces shaping it. Reasoning based on regular political/economic incentive structures (e.g. economic liberalism not wanting to reduce economic activity) will fly out the window at lightspeed once fear becomes a factor. I personally think its great that LLMs empower individuals such as your friend to increase the control they have over real issues in their life; that is what technology should be doing for us. Cloud inference-backed LLMs are doing the opposite: drastically reducing the power that the everyman has over his socio-economic future by throwing high-paying career paths for a whirl and incentivizing powerful organizations to destabilize the labor market. Only time will tell how this plays out.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#753

Earlier quoted context omitted.

Would really like to know how their watermarking technique works, and if they can use it to store arbitrary information in the text (I assume they can if only in a limited way). I assume if the text is long enough they could add all kinds of metadata that would then be undetectable as the model that generated the text is the "cryptographic" key that encodes the data. I wonder if you can have another model run over th…

I realize that short attention spans are pervasive now, but the link to the explanation is only eight paragraphs in https://declaude.org/watermarking/

In addition to holding the key, wouldn't you additionally need to know exactly which model to check against? So for passive detection to happen, I think each company would need to check every message against every model version? Also would need to spend resources re-invoking the each model version against each message.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#754
Get fucked? Really? As if the quality of machine-generated prose were somehow sacred?

Is Gruber only now waking up to the idea that LLM companies do not give a shit about the quality of the writing they generate? That they're disdainful of the entire concept of writing as a profession? What rock has he been under?

If you care one iota for the quality and craft of your writing, you would already recognize that any degree of AI "processing" obliterates stylistic choices. You're left with readable text, but it's fluffy, meandering, and has a cadence to the writing that screams to the reader "every second you spent on this was wasted."

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#755

You would think that if the EU is requiring this, then why not just implement this for the EU users.

Because it is not about the EU. Google has implemented watermarking for a long time. And soon we will see that all AI providers will do it, because it allows them to filter out AI generated data from their training sets.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#756

Earlier quoted context omitted.

In places like Germany it's a crime to say something that makes a politician look bad, even if it's true.

That is simply incorrect. In Germany, truth is a complete defense as far as libel and defamation cases are concerned.

But not insult, and even for defamation, you will be prosecuted and have to prove the statement is true, even if the politician and everyone else knows it's true.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#757

Earlier quoted context omitted.

Based on the SynthID-Text paper https://www.nature.com/articles/s41586-024-08025-4 I agree that the LLM's learned distribution isn't modified, but I don't think it's correct to say that the sampling process is not modified. Also I just read the paper today so I could be misinterpreting things. As described in the paper, you're right that it doesn't affect the main sampling technique, but what they do is they sample t…

I think you're correct. It does alter the distribution for each output token, implicitly giving each candidate token a different probability. Maybe that's fine, but it's not as magical as Anthropic [and a lot of commenters here] are making it out to be.

Update: https://x.com/mrcslws/status/2089850106292162982

Pasted below:

Finally convinced myself that non-distorting watermarking is real, a la Anthropic / Google SynthID. It's not just spin / marketing. This really is a "Monty Hall" like problem. (h/t @random_walker for that analogy.)

Sharing here in case it helps anyone else.

On one hand, watermarking is "obviously" damaging to the output. The LLM does all this work to compute token probabilities... then you essentially perturb the probabilities? Of course that's bad! Every single token is using perturbed probabilities!

On the other hand, you obviously can freely inject a layer of randomness into sampling. Take any categorical distribution. For each sample, perturb the probabilities, then draw the sample. Done correctly, over many samples the results will match the original distribution.

It still kinda feels like magic, but it's not surprising that you can take that core trick and shape it into a non-distortionary watermarking scheme.

(i.e. it gives you sequences that were just as likely to be generated by the original model as any other sequence)

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#758

> "The exact words we choose when writing matter." Then write your own damn text if you care about the exact wording so much

He's a professional writer. As the article says, even if you write your own words, this is still a problem with proofreading, copy-pasting references or quotes, and "AI checkers".

Is the concern that the AI will rephrase the text during proofreading, thus falsifying a copy-pasted quote? I'm not sure how such a change could go unnoticed, unless you blindly publish the AI output without looking at it, which doesn't sound like something someone would do who wants to “write their own words”.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#759
post #293

Earlier quoted context omitted.

I think this is a very real concern. But I’m not sure of any way around it. Any stenographic system that you have the code for can be trivially defeated. I wonder if this would be a good use for homeomorphic encryption. There might be a way to let anthropic check some text without actually giving them access to the source text. Any experts around? We could use your skills!

> Any stenographic system that you have the code for can be trivially defeated. They're giving you an oracle regardless, which is almost as good. Take LLM output, make some modification, ask the detector if it's LLM output, repeat until you learn what kind of changes you have to make to defeat it. Or don't even bother learning what to do, just make arbitrary changes until it says it's not, so when the person they're…

Good lord, just write the thing.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#760
post #9

> I want any LLM I use to choose the very best, most precise words at every single decision point. Then bad news: LLMs already use randomness in a fundamental way. Each time they go to generate a token, they first generate a probability distribution of possible tokens. Then they pick one randomly according to this distribution. The technique described can be thought of as making the random number generator pseudo ran…

The best discussion I've seen so far is from Scott Aaronson: https://scottaaronson.blog/?p=6823 > To illustrate, in the special case that GPT had a bunch of possible tokens that it judged equally probable, you could simply choose whichever token maximized g [a cryptographic function]. The choice would look uniformly random to someone who didn’t know the key, but someone who did know the key could later sum g over all…

> instead of selecting the next token randomly, the idea will be to select it pseudorandomly, using a cryptographic pseudorandom function, whose key is known only to OpenAI.

Seems like - given enough text to encode information into - it would be possible for OA to uniquely identify the user account (and maybe even the specific request) that generated some content, even if the chat text itself isn’t stored.

Interesting argument in favor of local AI as a mechanism for privacy-preserving generated content. Though, I wonder if there’s a way to “bake” a hardware fingerprint into local models as well…

Post reply on HN