Live data from Hacker News

How Claude marks AI-generated content

support.claude.com

191–200 of 446 posts

Re: How Claude marks AI-generated content

#191
post #88
post #15

> When a supported Claude model generates text, it weaves an imperceptible watermark directly into the text itself. You won’t see it, and it doesn’t change the meaning, quality, or readability of Claude’s response. I'd like to know a lot more about how that works. A lot of my interactions with Claude return pretty precise text. If I ask it to edit a project and refactor a specific function in several places I know ex…

It was quick :) … https://claudewatermarkremover.app/

> Honest note: Anthropic has not shipped a public Claude watermark detector yet. This tool uses rewrite-based neutralization — a meaning-preserving paraphrase with a non-Claude model — which is the attack path watermark research points to. Not affiliated with Anthropic.

Well, they should have run their own AI slop website through their tool...

Re: How Claude marks AI-generated content

#193

I notice the "Limitations" section talks about how content only at some point touched by Claude may return a positive, and content that returns a negative may still be Claude generated. But I really would have liked for them to state explicitly that entirely false positives where a piece is fully human-written may still be marked as generated, because too many institutions with the power to ruin someone's life over t…

If there is any false positive rate (which, because text will naturally and by chance include tokens from the green and red sets in some pattern, there will be), tools making promises like "detect AI-generated text" are unacceptable. They are going to turn innocent people into pariahs on some unsubstantiated "this content is 37% likely to be AI" claim that the user has no way of verifying or inspecting more deeply, we just have to trust the statistical box and assign some meaning to whatever that number means. 37% of my phrases are AI? There's a 37% chance my entire text is AI written? Part of the fun is not knowing!

This is scripture homeopathy and it's irresponsible.

Re: How Claude marks AI-generated content

#194
Why would I want a stochastic parrot that intentionally speaks less probable tokens?

Anyway, I've found a magic line that can be copied & pasted to the comment sections of most OpenAI/Anthropic news threads. This one is no difference.

The magic line:

> Doesn't matter; have DeepSeek.

Re: How Claude marks AI-generated content

#195
All big LLMs already visibly watermark all their text with easy to detect annoying phrases and turns of speech that everyone is already sick of hearing. Why do AI companies keep making their products worse to appease anti-AI, it’s not like they’ll suddenly start supporting it if you do so. If you’re worried about European customers, just relax your firewalls to let more VPNs through, if the productivity boost is high they will use it anyway if their rules keep crippling their own models

Re: How Claude marks AI-generated content

#196

People with dyslexia and dystrophia, commonly use LLMs to proofread content. Even Anthropic admits this is a limitation.

Yes, I’m audhd and dyslexic. I am cancelling my Claude max 5x subscription and moving to ChatGPT pro. I have difficulty enough trying to ensure my meaning comes through correctly, along with everything else; to now have to look out for/analyse watermarks too? I feel shamed enough by society, thanks Anthropic.

If it's only proofreading text you've written, the changes will be minimal enough that watermarking seems impossible to me.

Re: How Claude marks AI-generated content

#197
post #6

I have had a hunch for a while now that (in addition to these tools), Anthropic has actually leaned in to Claude's distinctive manner of writing since it makes the text more obviously AI generated and thus less susceptible to misuse. That's not necessarily the same thing as a markov-style fingerprint but it could be a correlated factor.

It's pretty trivial to command it to not speak that way. That's one of the first things you should write into the prompt. What style you want it to write in. Make it use a very concise and dry academic style with no overt LLMisms, melodramatic or flowery language, or metacommentary.

People have been posting some variant of this comment for three years, and it's no more true today. Ever notice that the "prompt engineer" career hasn't materialized?

Re: How Claude marks AI-generated content

#198

People with dyslexia and dystrophia, commonly use LLMs to proofread content. Even Anthropic admits this is a limitation.

Yes, I’m audhd and dyslexic. I am cancelling my Claude max 5x subscription and moving to ChatGPT pro. I have difficulty enough trying to ensure my meaning comes through correctly, along with everything else; to now have to look out for/analyse watermarks too? I feel shamed enough by society, thanks Anthropic.

How could "your meaning" come through if a computer is writing it?

Re: How Claude marks AI-generated content

#199
post #72

If the western AI companies are forced to comply with this type of BS, and develop their models to do their job while balancing a book on their head and hopping on one foot, the Chinese models just got a free pass to completely dominate the frontier. EU regulation does it again!

Less AI slop sounds like a win to me. Let China drown in it.

Re: How Claude marks AI-generated content

#200
post #15

> When a supported Claude model generates text, it weaves an imperceptible watermark directly into the text itself. You won’t see it, and it doesn’t change the meaning, quality, or readability of Claude’s response. I'd like to know a lot more about how that works. A lot of my interactions with Claude return pretty precise text. If I ask it to edit a project and refactor a specific function in several places I know ex…

>I'd like to know a lot more about how that works. My guess is that it works like Gemini's SynthID: by altering the logprobs of the next token. Like, for every 10th token, instead of outputting the most probable, it outputs the 17th most probable, or something. (Obviously it's way more complicated but I think conceptually this is how it works.) No human will notice this, but a classifier trained on Claude's output wi…

> Like, for every 10th token, instead of outputting the most probable, it outputs the 17th most probable, or something.

Wouldn't you need the prompt to know the probability of the next token?

Post reply on HN