Live data from Hacker News

Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

daringfireball.net

651–660 of 776 posts

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#651
post #625

Earlier quoted context omitted.

Easy to say, hard to come up with a believable alternative. In the meantime, you're flunking out a lot of people for having the integrity to not cheat and as a result not being able to keep up with an artificially inflated workload. You can't just destroy some signal and handwave that you'll make it up in some other way.

> the integrity There is none. It's just a jobs program fueled by student loans. Higher education in the west has been corrupt for quite a while now. AI is just the final nail in its coffin.

This is far too cynical. I use what I learned in my computer science program constantly.

I agree that higher education shouldn’t be as required to get a good job. But reality isn't black or white. Reducing the entire sector to a corrupt degree mill throws the baby out with the bath water.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#652

Earlier quoted context omitted.

> especially not for idiotic reasons like facillitating AI stigmatization. From the people I’ve talked to at universities, LLM based cheating in education is an unstoppable nightmare. I don’t have a problem with LLMs. But I do want the cheating to - somehow - stop. The people who cheat miss out on learning. And the people who don’t cheat have their degrees devalued by those who do cheat.

> The people who cheat miss out on learning. They aren't there to learn. They are there to jump through hoops to get a degree that will let them get a job so they can make money and prosper . The learning is entirely secondary. The cheating will stop when there is no longer any economic incentive to be there in the first place. People with "pure" motives will refuse to cheat on their own, precisely because they want…

Got any evidence for that claim?

I’ve worked with plenty of smart, self taught programmers throughout my career. The highest paid guy I know didn’t finish high school.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#653

Crazy how a smart person like this fails to understand the gumbel softmax technique. It does not affect writing quality at all, provably. The very fact that there is generally no "best next token" with 100% certainty is precisely why the trick works (you cannot watermark a response to "respond with the To be or not to be soliloquy from the first folio Hamlet", for precisely this reason).

> The very fact that there is generally no "best next token" with 100% certainty

This is not entirely accurate. Sure, there's never a token with 100% certainty, but there are often tokens with 99.9% probability, but this technique of course does not change how such a token is sampled.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#654

Earlier quoted context omitted.

And I'm sure students will use tools to have every other paragraph written in the style of a different AI, in an attempt to defeat this fingerprinting.

Yeah, wait for LLM "scrambles" that put every paragraph and then the whole text through multiple re-write/edit style cycles.

How would that change anything? The proposed watermark is applied while the output tokens are being chosen, taking that text and running it through an LLM again would just repeat the process.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#655

Earlier quoted context omitted.

> It is also in the AI companies' own best interest to be able to detect AI generated text so that they can avoid training on it It’s also in their interest to demonstrate that they can be trusted and to show that they at least pay lip service to limit the obvious downsides of the tools they are selling. The use cases they sell to mainstream audiences are not affected by detection tools. The point of having a LLM do…

The EU mandate for AI watermarking is likely the first step in the direction of prohibiting AI for specific use cases. The pretext for outlawing (or at the very least controlling) the use of AI when the time comes will be something along the lines of data integrity or just general compliance legalese. So, the use cases that are being sold to mainstream audiences actually will be affected by detection tools, especiall…

I’m sorry, this is going to be a bit long but you made good points.

> The EU mandate for AI watermarking is likely the first step in the direction of prohibiting AI for specific use cases.

I am not sure how practical that would be. The cat’s already out of the bag and they won’t prevent companies in the whole world from releasing open weight models. Playing catch up by distilling flagship models is also relatively cheap; we’d see smaller companies setting up shop in friendly regimes. And I don’t see any appetite to go full child porn and criminalise the possession of a LLM. So we’d end up with a similar situation as with illegal downloads, i.e., everyone will do it and nobody will care.

> The pretext for outlawing (or at the very least controlling) the use of AI when the time comes will be something along the lines of data integrity or just general compliance legalese.

They could forbid using LLM for hacking, but hacking is already illegal. They could make it a factor when determining punishment, but I don’t think that would work terribly well. Most of the dangerous stuff we can do with LLMs is already illegal, or should become so. Things like propaganda, identity theft, harassment, scams. We need enforcement with teeth on these, not pointless feel-good legislation. Again, there are parallels with cryptocurrencies and torrenting software. These things have illegal uses, but it’s also really difficult to make them illegal, at least in semi-functioning democracies.

> So, the use cases that are being sold to mainstream audiences actually will be affected by detection tools, especially if the output is intended to be monetized in some way.

I don’t know that mainstream audiences are really against LLMs. They are mostly against AI in a nebulous sense, but even non-technical people use ChatGPT or equivalent. I think that the critical mass is already there and the tools are convenient enough that they couldn’t outlaw them without a massive uproar.

Detection tools don’t seem all that relevant to mainstream audiences’ use of LLMs. AI companies will sell this as a safeguard against misuse, and everyone will be happy about it. The politicians will say they accomplished something, the AI companies will slowly turn public opinion, and the public will have shiny toys.

> The number of legitimate use cases for costly frontier models drops significantly once eliminating professional jobs is off the table.

I don’t know. They can open possibilities that we don’t necessarily consider.

One example I have is a friend who is getting his house refurbished. He’s not an engineer or a material scientist. He does not have enough free time to read thoroughly on the many subjects involved. With a decent LLM, he could untangle the technical documents sent by the architect and the contractors to really understand what was going on and be involved, rather than passively follow the architect’s advice. For starters, the LLM was very useful in finding issues in the quotes he received when he was looking for an architect. Those were long, technical documents, with no really standardised structure and full of jargon. I don’t think that person is going to want to stop using LLMs now. Many people are having this sort of moments right now.

> In the near future the EU will likely come down with heavy intervention to prevent AI from impacting employment rates across Europe.

Maybe. But i don’t believe the EU is well equipped for that. Labour laws are largely local and different in each member state. The EU regulations are basically the common denominator, ore or less, and it is easy to see why: for regulations to get adopted, they need a strong enough majority in the Commission, in the Parliament, and in the Council. It is very difficult to get anything controversial that affect the sovereignty of member states passed.

The angle of the current AI regulations is that they set the rules for the single market, which is where the EU is the most legitimate. It is difficult to see market angle for the effect of AI in labour, and I think enough member states would be keen to kill the project.

Also, there are many influences at play, but the EU is fundamentally an economically liberal institution. It very rarely goes in the direction that reduces economic activity. Look at how clumsy it is at fighting against cheap Chinese imports. I don’t think the institutions themselves would really want to make AI illegal. Companies have too much to lose.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#656
post #246

My biggest concern is that checking any text for watermarks requires sending the entire text to Anthropic. And even that is not sufficient, as the text might have been generated with ChatGPT, Gemini, Grok, Mistral, ... So every check requires sending the text to as many AI providers as offer a watermarking detection API, almost all of which have a very dubious track history with obtaining training data through illici…

Exactly, so even if you are avoiding AI, any interaction with society is now being structured so you have to submit to the digital surveillance equivalent of a cavity-search machine. What a dystopia awaits the budding generations.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#657

Earlier quoted context omitted.

I am sure he is quite happy to have people disagree with him on HN. He usually wears it as a badge of pride. We might even have a follow-up in a couple of days about how these techy weirdos lost the plot.

Indeed.

:D

(For the record, I read DF quite often as a Mac-minded techy weirdo, I just happen to disagree on this particular issue)

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#659

Earlier quoted context omitted.

huh, thanks for commenting this! I trusted the linked explanation post https://declaude.org/watermarking/ but actually reading the the synthid paper showed me that my understanding was wrong: https://www.nature.com/articles/s41586-024-08025-4 It is definitely blurrier whether you can say this approach changes the distribution then. By definition, it _has_ to change the probabilities of output tokens, but it's not tot…

You’re not getting it. The probabilities do not change at all. The only change is given some probabilities there is a deterministic method for determining which symbol was sampled from that distribution. The distribution or sampling process itself is not modified.

Based on the SynthID-Text paper https://www.nature.com/articles/s41586-024-08025-4 I agree that the LLM's learned distribution isn't modified, but I don't think it's correct to say that the sampling process is not modified. Also I just read the paper today so I could be misinterpreting things.

As described in the paper, you're right that it doesn't affect the main sampling technique, but what they do is they sample the distribution for 2^m samples, and then use Tournament sampling to choose the tokens among those 2^m samples, and the watermark key changes the scoring of the tournament options, using the watermark key as an input to the random generator that generates the scoring functions.

Then, to calculate the watermark, they take the text, and compute the mean g-values of the text, and a higher score means that it's more likely that it was sampled using the provided selection of tournament watermarking functions.

let's say you had some top P words: mango, banana, pineapple, guava, and you sampled 8 times, and got each one twice in the following order:

1. mango 2. banana 3. pineapple 4. guava 5. mango 6. banana 7. pineapple 8. guava

without tournament sampling, you'd truly see any of those come through. But in tournament sampling, you take those 8 options, create m scoring functions based on the pseudorandom generator, and score the 'tournament' by sampling the biased distribution you create from the g values. That does change the sampling from based purely on the LLM and entropy, but i mean, if the watermark key is also generated from some entropy, it's probably representative of the original sampling options as expected?

this is a very fascinating topic! I do still stand by my point that anthropic is the only one who can tell if something is watermarked or not and feeling icky, but the paper has mostly quelled my concern on impacting the intelligence part.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#660

Earlier quoted context omitted.

I’m sure it matters. But how much does it matter? How much (perceived) intelligence would you be willing to sacrifice for an accurate AI predictor? I’d sacrifice a few %, easily. Maybe 10%. The models are getting smarter at such a fast rate that I’d be willing to lose a month or two of progress to help slow down the AI cheating epidemic. It sounds like you expect this fingerprinting approach would dramatically reduce…

10% is a lot! And why would this slow down the cheating epidemic? There’s tons of ways like declaude etc to get around the check. Also Claude’s style is very distinctive (e.g. “load bearing”) AND has changed since 4.5 quite dramatically. I might not be able to tell on a specific response, but I can tell you that I went from canceling my ChatGPT plan in November, to now reaching for it first and considering cancelling…

> There’s tons of ways like declaude etc to get around the check

How effective is this on the new fingerprinting mechanisms?

> Is the cheating epidemic so bad?

From what I’ve heard, yeah it’s out of control. And all the existing llm detectors that academics use have a high false positive rate, which catches a bunch of innocent students in the cross fire.

> it does feel like the assignment and ways education happens needs to change?

Why? Was there something fundamentally wrong with how universities teach and assess?

The sector is responding. For example by moving back to more in person exams and reducing the load of any take home exams. Is that good, for some reason?

> There’s something Orwellian about how the phrasing of a passage embeds hidden information

Interesting. I don’t have the same response. LLMs give me an acute sense of existential dread each time their capabilities improve. But fingerprinting doesn’t move me at all. Do some soul searching on why this bothers you. I’d love to hear why, and I bet you aren’t alone.

Post reply on HN