Live data from Hacker News

How Claude marks AI-generated content

support.claude.com

211–220 of 446 posts

Re: How Claude marks AI-generated content

#211

Earlier quoted context omitted.

That's essentially impossible, unless you mean they didn't measure a false positive rate.

For watermarked long-form text, it is actually possible. Makes the watermark more fragile, but the math is considerably more forgiving than usual.

And yet, it remains possible that a human could write the same sequence of characters.

Re: How Claude marks AI-generated content

#212

We need to just stop pretending we can reliably tell if plain text is written by an LLM. It’s just not a reasonable ask.

True, but what you can do is a one-sided guarantee. If it bears the mark, it is likely generated (or someone deliberately made it look generated). Thus, if a news article, research article, book, student paper submission, blog post , HN comment, etc, bears the mark, it could be automatically flagged as such. It helps detect low effort slop. --- Caveat. If you write your own creative work and send it to Claude for "cl…

Not might, will. Whether enough text is present or not to go over the detection threshold is in doubt. But the "score" will never be zero, even for human written text.

Re: How Claude marks AI-generated content

#213
If they can use a cryptographic key to sign the text, will they generated keys unique to users? Seems like they'd be able to trace back to which accounts were used to say generate scam dialogue, threats, propaganda or bot content on the open internet?

Re: How Claude marks AI-generated content

#214

I notice the "Limitations" section talks about how content only at some point touched by Claude may return a positive, and content that returns a negative may still be Claude generated. But I really would have liked for them to state explicitly that entirely false positives where a piece is fully human-written may still be marked as generated, because too many institutions with the power to ruin someone's life over t…

Maybe it has no false positive rate

Somewhat trivially, if I ask Claude to transcribe an image and then check if that transcription is ai generated it will likely say yes.

Many users are not smart enough to realize that the transcription step is where the ai (watermarks) were necessarily injected.

Re: How Claude marks AI-generated content

#215

I notice the "Limitations" section talks about how content only at some point touched by Claude may return a positive, and content that returns a negative may still be Claude generated. But I really would have liked for them to state explicitly that entirely false positives where a piece is fully human-written may still be marked as generated, because too many institutions with the power to ruin someone's life over t…

If there is any false positive rate (which, because text will naturally and by chance include tokens from the green and red sets in some pattern, there will be), tools making promises like "detect AI-generated text" are unacceptable. They are going to turn innocent people into pariahs on some unsubstantiated "this content is 37% likely to be AI" claim that the user has no way of verifying or inspecting more deeply, w…

I'm curious about your thoughts on pangram. I only really see posts on Reddit claiming it falsely labels their content as ai generated but nobody will actually post examples of "textbook from twenty years ago" or upload screenshots of a journal (also those posts usually feel deeply ai generated without an ai detector)

Do you think this is an impossible task and we shouldn't try to solve it? Or do you think it's doable and that some ai detectors might be better than others?

Re: How Claude marks AI-generated content

#217
> When a supported Claude model generates text, it weaves an imperceptible watermark directly into the text itself. You won’t see it, and it doesn’t change the meaning, quality, or readability of Claude’s response.

I feel like this is FUD. If you copy text from Claude, Ctrl Shift V it into VS Code, the IDE will give up the ghost on if weird characters are in there. And it's not like Google suddenly invented new letters or fonts either.

Practically speaking, I feel this is Google publishing misinformation.

Re: How Claude marks AI-generated content

#219

Earlier quoted context omitted.

It's pretty trivial to command it to not speak that way. That's one of the first things you should write into the prompt. What style you want it to write in. Make it use a very concise and dry academic style with no overt LLMisms, melodramatic or flowery language, or metacommentary.

People have been posting some variant of this comment for three years, and it's no more true today. Ever notice that the "prompt engineer" career hasn't materialized?

Regular engineers still exist.

Re: How Claude marks AI-generated content

#220

People with dyslexia and dystrophia, commonly use LLMs to proofread content. Even Anthropic admits this is a limitation.

Yes, I’m audhd and dyslexic. I am cancelling my Claude max 5x subscription and moving to ChatGPT pro. I have difficulty enough trying to ensure my meaning comes through correctly, along with everything else; to now have to look out for/analyse watermarks too? I feel shamed enough by society, thanks Anthropic.

Have you tried Kagi Proofread?

https://translate.kagi.com/proofread

Post reply on HN