Live data from Hacker News

Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

daringfireball.net

401–410 of 776 posts

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#401

Text watermarking is another EU rule made without real world input. The Union is stuck on major economic crises (electricity prices for instance) because nobody can agree on anything. However, the bureaucracy forces tech into a privacy nightmare. Brussels cannot bring together its own members but it loves pretending it can govern the internet.

> (electricity prices for instance) because nobody can agree on anything.

I would say that's more like because the US has arranged for Europe's fossil fuel energy sources to be disrupted or cut off:

* Libya - NATO made a pig's breakfast of that, it's a failed state now.

* Iran - transitive sanctions, because why not prevent non-US states from trading with each other.

* Russia (& Kazahkhstan) - The US (with or without Ukranian involvement) bombed the NordStream pipeline(s), led the EU into the proxy war in Ukraine and a sanctions regime against Russia. Kazakh oil goes to Europe through Russia.

* Gulf states - until recently, possible but not very convenient ; since Feburary of this year, the war on Iran messed that up badly too.

the US is the winner here not just geo-politically, but also as an oil exporter, with the EU now depending on purchasing US-exported oil.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#402
post #32
post #16

This is not meant to be snarky, But almost any writing done by Claude is a perversion of writing. I honestly can't stand the way Claude writes. This watermark change just makes it scarier.

I moved to Sol for my writing and it is so so much better. But it makes more mistakes. I think they have different ideas of product but it seems OpenAI is going to follow Anthropic’s lead over the next year. I think I am going to put more effort into my writing skills to remove myself from this awful situation

> I moved to Sol for my writing and it is so so much better.

No it's not. The bad part about it is that some machine is writing instead of you, not the specific stylistic idiosyncracies.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#403
post #246

My biggest concern is that checking any text for watermarks requires sending the entire text to Anthropic. And even that is not sufficient, as the text might have been generated with ChatGPT, Gemini, Grok, Mistral, ... So every check requires sending the text to as many AI providers as offer a watermarking detection API, almost all of which have a very dubious track history with obtaining training data through illici…

its lovely training data. no detection? add to training set -_-. its also kind of laughable that somehow people are trying to prevent the outputs not to be altered. Asif you cannot manually paraphrase anything you can read. So the only solution would be, to make it utterly unreadable (which is not possible, it obviously defeats the purpose of the thing). Not to mention local models ofcourse :-)

I think what we’re testing for here is LPM output that hasn’t even been skimmed by a human, let alone paraphrased.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#404
post #377
post #362

Earlier quoted context omitted.

> Why not switch it around? Because a malicious human will gladly copy/paste LLM text and sign it with his "I, a human, definitely wrote this academic paper" key?

Fair point, for text it is far harder to prevent signatures being applied to generated text vs images at the moment of capture and most approaches I can come up with to remedy this can either be bypassed (edit histories can be output by models similar to humans) or will be controversial. Taking a page out of the anti-cheat textbook, mainly written for gaming, there are methods which might hold in the medium term. Les…

I've been thinking for a while that all of this is just trying to grasp tighter the last bits of sand escaping between our fingers. The end game, perhaps, is trust. Do you trust or know the source? If you don't, assume it was AI generated. If you do, accept it as authentic based on whatever they disclose, but know that it's possible they aren't being totally honest or were themselves fooled in some way, depending on the context.

Then build our assumptions and how we operate around those trust levels in the digital realm.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#405

Earlier quoted context omitted.

You know, back in the era when proofreaders were human, I never one met a proofreader who rewrote my text afresh, rather than annotating the text with a pen. It's still possible to use Claude to proofread - highlight grammatical, flow, structure, logic errors and make simple suggestions for you to pick and choose or adapt as you wish. No watermarking will flag your text. No flaw accusations of LLM authorship will hau…

Where is the problem with using LLM generated text? You could use your own hypothetical house elf to do it for you, or pay someone to do it. LLMs are just cheaper for a certain set of problems. People will find ways to circumvent this, so this limitation will only hit the technically less adept people.

If there was a ghostwriting detector, I would use that too.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#406

Earlier quoted context omitted.

> It seems to me like he started out mad and looked to justify it. Yes, but that's neither surprising nor a reason to dismiss the anger. People get angry about DRM schemes in video games, even if the slowdown these cause is practically imperceptible. They're angry -- and Gruber acknowledges that factor too -- because a stranger manipulates what they regard as their own domain, without consent by or benefit to the own…

> People get angry about DRM schemes Yes but thats a thing that degrades something in a catastrophic way, as in I can use the thing one day, and not the next. A different randomisation system on something that is a text generator which is designed to be unperceptable sounds like the people who are annoyed at FLAC vs MP3[1] Done right you won't know the difference, done badly and you will. [1] ex audio engineer, try m…

> the people who are annoyed at FLAC vs MP3

What’s the story there? I didn’t know that was a thing and I’m curious to learn more.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#407

Earlier quoted context omitted.

There are diffusion-based models and transformer-based models (and many other "architectures"), so your comment does not make sense.

Are there any diffusion-based or otherwise non-transformer-based models in mainstream use?

If you include non-language models, yes.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#408
post #406

Earlier quoted context omitted.

> People get angry about DRM schemes Yes but thats a thing that degrades something in a catastrophic way, as in I can use the thing one day, and not the next. A different randomisation system on something that is a text generator which is designed to be unperceptable sounds like the people who are annoyed at FLAC vs MP3[1] Done right you won't know the difference, done badly and you will. [1] ex audio engineer, try m…

> the people who are annoyed at FLAC vs MP3 What’s the story there? I didn’t know that was a thing and I’m curious to learn more.

MP3 is lossy. FLAC is lossless. So obviously a certain type of people are going to make a religious war out of it.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#409
post #246

My biggest concern is that checking any text for watermarks requires sending the entire text to Anthropic. And even that is not sufficient, as the text might have been generated with ChatGPT, Gemini, Grok, Mistral, ... So every check requires sending the text to as many AI providers as offer a watermarking detection API, almost all of which have a very dubious track history with obtaining training data through illici…

Doesn’t this mean Anthropic can accuse anyone of using their AI to write for them?

Your favorite anti-AI political candidate turns out to have not written their thesis, with a 73% confidence level.

Re: Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

#410
post #396

Earlier quoted context omitted.

My main point is that sampling with a modified distribution compared to the one produced by the model is already being done, and it is generally found to increase quality, not decrease it. So there is no reason a priori to assume that the watermarked distribution would be lower quality than other schemes for altering the "raw" output distribution (such as top P, top K, temperature, etc). My second point is that the t…

All those methods are applied with the specific goal of improving output quality and are applied to the extent that they do this. Watermarking has no such goal, and is not implemented for any such reason. In fact it's much more like applying another layer of random noise over the token selection process, because the sequence that generated the green token list comes from a seeded PRNG. >My second point is that the tr…

> All those methods are applied with the specific goal of improving output quality and are applied to the extent that they do this

Yes, that's the goal that was used, but they are quite simplistic and crude methods, not some specifically designed function, with carefully fine tuned parameters or something. So, if a basic function like top_k can improve model utility, it's not impossible to imagine that watermarking could also happen to do so, or at least not have a significant negative effect. So whether the effect is deleterious or not is an empirical question, not something we can assume ahead of time.

> You're basically saying that if we train LLMs on watermarked text they'll be really good at producing text that looks watermarked

No, you're misunderstanding how the training works. If we train the model's output so that it minimizes the error function after the watermark is applied on it, the model will learn how to produce the best output it can given the watermark. It will produce better text that happens to be watermarked, not "more watermarked text". Same as if you train the model on minimizing `top_k_error(input) = |top_k_sampling(model_output(input)) - desired_output(input)|`, the model will learn to produce better output under top_k sampling, not learn to produce output that's "looks more top_k".

Post reply on HN