Live data from Hacker News

Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

daringfireball.net

491–500 of 776 posts

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#491

Earlier quoted context omitted.

1) we’re not discussing those systems. We’re discussing a chat AI product called Claude, which does not offer those knobs. 2) Claude’s PRNG having a P is immaterial

Claude has those knobs, they are just not exposed to the user. They could make Claude nearly completely deterministic if they wanted to (of course it would be a far inferior product then. But they could).

“Not exposed” = has no knobs. Of course all autoregressive LLMs can be operated this way but Claude, the product, employs LLMs but isn’t one.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#492
post #246

My biggest concern is that checking any text for watermarks requires sending the entire text to Anthropic. And even that is not sufficient, as the text might have been generated with ChatGPT, Gemini, Grok, Mistral, ... So every check requires sending the text to as many AI providers as offer a watermarking detection API, almost all of which have a very dubious track history with obtaining training data through illici…

That's just storing what they output in a database and then checking, not a watermark.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#493
Wait are we supposed to be mad because clankers are displacing human creative workers, or mad because clankers don't do their level best when producing creative works because they are forced to watermark text? Or is it that they use all the water (I know they don't but are we supposed to be mad about it still)?

I can't keep up with the current Chinese psycop. There should be some kind of status page like whywewantamericatofailataitoday.ai so we can keep up with it.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#494

> "The exact words we choose when writing matter." Then write your own damn text if you care about the exact wording so much

What an absurd take on a very reasonable concern. As much as I hate “obviously AI” writing, I don’t understand why we have to handicap the tech and prevent it from improving. This manipulation of word choices virtually guarantees AI will always sound like a robot.

> This manipulation of word choices virtually guarantees AI will always sound like a robot.

You make it sound as if that were a bad thing.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#495

Earlier quoted context omitted.

In practice, it is theater. Are they going to do this with the code output too? This is nonsense security theater for the low thinkers to have a sense that someone is in charge. When we all know nobody is in charge, anywhere.

EU regulations are going to force it, friend. The tsunami is coming and cannot be stopped and Anthropic has jack to do with it.

I wonder if that is entirely true. It is also in the AI companies' own best interest to be able to detect AI generated text so that they can avoid training on it ("Habsburg AI").

In the case of Anthropic it would also be entirely unsurprising if they've been lobbying the government to force everyone to do something in their (Anthropic's) own best interest.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#496
post #246

My biggest concern is that checking any text for watermarks requires sending the entire text to Anthropic. And even that is not sufficient, as the text might have been generated with ChatGPT, Gemini, Grok, Mistral, ... So every check requires sending the text to as many AI providers as offer a watermarking detection API, almost all of which have a very dubious track history with obtaining training data through illici…

> And even that is not sufficient, as the text might have been generated with ChatGPT, Gemini, Grok, MistraL

I would think that these things would eventually converge and we’d get one watermarking algorithm as an industry standard. That way, all major provider would follow it and we’d get independent software for checking. This would partly limit the efficacy of the watermarks, but on the other hand if it’s done correctly, removing the mark could still be enough of a pain that casual users would not bother. That would obviously depend on a lot of factors. It would at least add significant friction in the production of daily slop.

Of course it wouldn’t do much for thing like foreign propaganda but that’s a whole other discussion we need to be having.

> Any university using AI detection in their submission pipeline, or lawyers, editorialists, proofreaders that check for AI marks will be sending significant amounts of text like unpublished research, books, potentially internal documents and more, most of which is high quality human written, to dozens of AI companies, blindly trusting they won't train on any of that.

Isn’t it already what they are doing right now with some of the plagiarism detection tools? Not every university is going to have a representative corpus, and yet they are all using the software. So I guess the provider is doing the work of feeding all that data to their algorithm.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#497
No, machines being used to replace human writing is a perversion of writing. Some might call it worse than cannibalism[0]. If you're only noticing now because Anthropic is changing things behind your back, well... I've got some bad news for you, but the entire cloud-hosted subset of the AI space, especially Anthropic, is premised on the fact that doing things behind your back to their models is socially preferable, or worse, should be outright mandated.

While I generally hate legal mandates to stab your customers in the back, in my opinion there is no harmless way to use AI and mandatory text watermarking is a good bare minimum. The EU probably made the right call. The entire AI space - open models included - is predicated upon worker exploitation, replacement, and deskilling; we should at least be able to know how much of our media diet has Anthropic's fingerprints on it.

A lot of hay is made over the pretraining process in which copious amounts of stolen data are trained on; but parallel to this is a huge data labeling and human feedback operation staffed almost entirely by people in third-world countries with robust English as a Second Language (ESL) programs. The thing is, AI models already watermark their text, they just happen to do so with the textual watermarks of the Indians and Nigerians that the AI companies hired to do RLHF because they were cheap. That's why certain AI models love the word "delve" so damned much. It's neocolonialism, designed specifically to do the kind of replacement the anti-immigrant idiots keep screaming their heads off about[1].

Furthermore, as we've seen with Hank Green, even non-cannibalism-adjacent AI usage is a recipe for worker deskilling and AI psychosis. The other half of the RLHF pipeline is to turn a pile of compressed text into a chatbot that feeds you a steady drip of unsourced information while praising you every time you spot one of its lies and never saying no[2]. This is a recipe for addicting your customers.

Also, this might just be because this is on daringfireball.net, but I can't help but think the author has an axe to grind against the EU because the EU mandated Apple sign third-party app stores. The fact that he's balking at Anthropic not going along with gating the watermarks to just the EU seems downstream of this - "why aren't you maximally attempting to resist the EU?"

[0] https://www.youtube.com/watch?v=YCPAIg7RUq8

[1] To be clear, upwards of none of the far-right have actually clued into the fact that AI is trained by underpaid immigrants, mainly because it doesn't fit the narratives the people running the far-right want to push. There are some AI robotics companies that are even very explicit that their robots are there primarily to launder foreign labor into rich companies and make permanent labor arbitrage.

[2] Continuing on from [1], the far-right actually really loves AI specifically because it rarely says no to even stupid ideas, even if it's also a manifestation of everything they claim to hate.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#498

> "The exact words we choose when writing matter." Then write your own damn text if you care about the exact wording so much

Do people not realize this will apply to ALL Claude output, not just writing you ask it to produce? > Marks will apply to output from supported Claude models across Claude Platform (API), Claude, Claude Code, Claude Cowork, and Claude Tag, and wherever Claude is offered. I've never asked an LLM to generate writing I wish to post as my own. I don't understand why we think it is a good idea to fudge all output just so…

[deleted]

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#499
post #246

My biggest concern is that checking any text for watermarks requires sending the entire text to Anthropic. And even that is not sufficient, as the text might have been generated with ChatGPT, Gemini, Grok, Mistral, ... So every check requires sending the text to as many AI providers as offer a watermarking detection API, almost all of which have a very dubious track history with obtaining training data through illici…

Doesn’t this mean Anthropic can accuse anyone of using their AI to write for them?

yes, why would someone will use this tool for writing then?

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#500
post #220
post #196

The same absolute morons who gave us cookie consent strike again. I swear, one of those days I will get into politics just to fight those two things, and the cottage industry of batshit crazy lawyers that gave birth to those things.

The cookie consent banner is not the EU's fault. It's either don't track or ask for consent. The fact that the industry chooses to track is not on the EU.

Its more nuanced than that. Even companies that do not track, and use only essential cookies, ask for consent, as the consensus among compliance teams and external lawyers is "its safer this way". Thats the reality which the regulators failed to anticipate.

Of course this misses a bigger point that tracking in the web moved in a direction that requires no cookies whatsoever, and if anything, feels more pervasive than it ever was.

And it misses the even bigger point, that the morons who legislated cookie consent did not notice either of those two realities. And the same thing is already true with the AI act; the text watermarking is trivially defeated and everyone knows it. And I'd bet it will remain a requirement for the next decade or three.

Post reply on HN