Live data from Hacker News

How Claude marks AI-generated content

support.claude.com

301–310 of 446 posts

Re: How Claude marks AI-generated content

#301
post #237
post #214

Earlier quoted context omitted.

Somewhat trivially, if I ask Claude to transcribe an image and then check if that transcription is ai generated it will likely say yes. Many users are not smart enough to realize that the transcription step is where the ai (watermarks) were necessarily injected.

How is a perfect transcription of an image watermarked?

"Perfect as far as human perception can tell" is a weaker standard than "bit-to-bit copy." Maybe it's that?

Re: How Claude marks AI-generated content

#303

Many good comments here. It’s somewhat common for me to voice record say a blog post of product updates, more like a ramble. Then have Claude clean it up. Then argue back and forth about certain things until it’s good, then make a final pass sometimes to change a few key words. This is incredibly different than pure ai text. Presumably it will show as ai generated here, even though I would argue it is not really. So…

Perhaps. But it seems like your beef is not with the presence of watermarking, it's with what people will use that watermarking for. You're not directly harmed by that blog post being labeled as AI generated. In a hypothetical (but unfortunately likely) world where everything has passed through an AI's digestive system, nobody would care.

In the meantime, it is true that this takes something away from you. But it's something you were only recently given. Now you're not given quite as much, but readers are given a little more (or rather, there's less being taken from us!)

> This is incredibly different than pure ai text.

Ok. But it's still incredibly different from pure human text. I guess the question is which provides more value? Providing the information "this text is AI watermarked" to readers? Or allowing creators to lie and claim that AI processed text was 100% human generated? I agree that people assuming that "has AI watermark" == "is AI slop" is incorrect and causes some amount of harm, but having the watermarks also pushes back on a large amount of harm already being done.

(Personally, I'm skeptical that these watermarks will ever hold up to adversarial attacks, and they haven't claimed that they will. So I think it's the usual "casual liars will be caught, determined liars will get an additional thin veneer of respectability".)

Re: How Claude marks AI-generated content

#304
post #88
post #15

> When a supported Claude model generates text, it weaves an imperceptible watermark directly into the text itself. You won’t see it, and it doesn’t change the meaning, quality, or readability of Claude’s response. I'd like to know a lot more about how that works. A lot of my interactions with Claude return pretty precise text. If I ask it to edit a project and refactor a specific function in several places I know ex…

It was quick :) … https://claudewatermarkremover.app/

Well, if the model that page uses also falls under the EU act, the output will just be watermarked differently. ;)

Re: How Claude marks AI-generated content

#305

Earlier quoted context omitted.

> Like, for every 10th token, instead of outputting the most probable, it outputs the 17th most probable, or something. Wouldn't you need the prompt to know the probability of the next token?

Not necessarily. Here's a rough example (it's not what's going on here, just a representative idea): There are words/tokens that are heavily correlated to the prompt (a yes or a no, for example), and then there are others that are going to be less so (adjectives with a lot of synonyms for example). Given a text, you can identify what the "load bearing" and auxiliary words/chunks are. Then, looking only at the auxilia…

So this is finally what "load bearing seam" means.

Re: How Claude marks AI-generated content

#306

I guess this is where our true colors show. There's a significant contingent of HNers who always dunk on LLM text detectors and claim that they can't possibly work, that they ruin careers, etc. But now that a lab says "OK, we'll add a real watermark", the reactions are overwhelmingly that it's still somehow wrong. Why do feel so entitled to being able to pass LLM-generated text as our own? I get that a lot of techies…

> Just because we found a "cheat" button doesn't mean it's wrong for others to want to know.

One difference perhaps is that you think using LLMs is cheating, while others do not.

Re: How Claude marks AI-generated content

#307

I guess this is where our true colors show. There's a significant contingent of HNers who always dunk on LLM text detectors and claim that they can't possibly work, that they ruin careers, etc. But now that a lab says "OK, we'll add a real watermark", the reactions are overwhelmingly that it's still somehow wrong. Why do feel so entitled to being able to pass LLM-generated text as our own? I get that a lot of techies…

It saves the non-anglophones the bother of learning to write readable English, so there's that.

Re: How Claude marks AI-generated content

#308

Like others have said, it's not reasonable to ask this. I propose we defeat this with the obvious: Simply, figure out what are some of the markers Claude and others will use for these tools, and sprinkle them randomly on everything we type or produce, all the time, 100%. If users flood the tools, and everything returns as AI-generated, then the tools become useless.

Why would you want AI content to be indistinguishible from human output?

Re: How Claude marks AI-generated content

#310
post #15

> When a supported Claude model generates text, it weaves an imperceptible watermark directly into the text itself. You won’t see it, and it doesn’t change the meaning, quality, or readability of Claude’s response. I'd like to know a lot more about how that works. A lot of my interactions with Claude return pretty precise text. If I ask it to edit a project and refactor a specific function in several places I know ex…

Certainly it can't watermark text with low entropy. If you're renaming a function using claude it won't be marked
Post reply on HN