Live data from Hacker News

How Claude marks AI-generated content

support.claude.com

111–120 of 446 posts

Re: How Claude marks AI-generated content

#111
post #29
post #15

> When a supported Claude model generates text, it weaves an imperceptible watermark directly into the text itself. You won’t see it, and it doesn’t change the meaning, quality, or readability of Claude’s response. I'd like to know a lot more about how that works. A lot of my interactions with Claude return pretty precise text. If I ask it to edit a project and refactor a specific function in several places I know ex…

> I'd like to know a lot more about how that works. Count load-bearing words using two different algorithms in a belt-and-braces fashion

One thing worth flagging: those words are load-bearing

Re: How Claude marks AI-generated content

#112
post #72

If the western AI companies are forced to comply with this type of BS, and develop their models to do their job while balancing a book on their head and hopping on one foot, the Chinese models just got a free pass to completely dominate the frontier. EU regulation does it again!

If the Chinese want to sell to EU customers, they probably have to do the same.

Re: How Claude marks AI-generated content

#113
post #64

Earlier quoted context omitted.

Panagram is a scam.

(This is the part where you provide extensive extraordinary evidence to your claim)

Scam might be too strong a word but it certainly has far higher false positive rates than they are claiming, and their output is at best misleadingly presented: https://freddiedeboer.substack.com/p/i-wouldnt-say-pangram-i...

Re: How Claude marks AI-generated content

#114

Earlier quoted context omitted.

They still need to choose when to do that though. When I prompt the program to e.g. alter a bash script in a specific way or to recite a longer known text it can't go round and randomly exchange tokens. It has to somehow define what is a simple repeated text from a different origin and what is a novel generation.

I am wondering how that applies to newly generated code. Odd variable naming? Stylistic choices that are watermarked? Or as someone else noted further down in the comments, it could be more subtle: Between the first and second most likely choice, in certain positions it will consistently choose in a certain way.

> Odd variable naming? Stylistic choices that are watermarked?

Whatever it is, I'm sure it's load-bearing.

Re: How Claude marks AI-generated content

#115
post #61
post #33

Earlier quoted context omitted.

You forgot about the cases where (1) people don't care, (2) people want to say "I used an LLM for this". I'm convinced that those cases happen more often than you think. Why not cover them with a simple mechanism? It's also in the interest of AI companies who don't want to train on AI output.

Depends on the pushback in different sets of users. Students for example would clean it up.

Sure, but let's first find out how many % of people are willing to be frank about their AI usage, and/or don't care about it. My guess is it is worthwhile to do this.

Re: How Claude marks AI-generated content

#116
post #17

Earlier quoted context omitted.

I had a similar thought but I assumed they leaned in because it improved performance on coding or something like that

It could also partly be a byproduct of examples of claude writing being in the dataset, which of course anthropic has lots and lots of and they do train on.

no way. there's just no good excuse for why "load-bearing" and "worth flagging" are everywhere now, I've pretty much never seen that in the wild before

Re: How Claude marks AI-generated content

#117

People with dyslexia and dystrophia, commonly use LLMs to proofread content. Even Anthropic admits this is a limitation.

How exactly does this impact proofreading? You can manually apply the suggestions (typo here, unnatural sounding sentence there, etc.) the LLM gives you to your own content, and it would stay watermark-free.

Unless with "proofreading" you actually mean having the LLM write your content for you.

Re: How Claude marks AI-generated content

#118

Earlier quoted context omitted.

I am wondering how that applies to newly generated code. Odd variable naming? Stylistic choices that are watermarked? Or as someone else noted further down in the comments, it could be more subtle: Between the first and second most likely choice, in certain positions it will consistently choose in a certain way.

> Odd variable naming? Stylistic choices that are watermarked? Whatever it is, I'm sure it's load-bearing.

You're absolutely right. But it is not just load-bearing, it is the load-bearing seams.

Re: How Claude marks AI-generated content

#119
post #29
post #15

> When a supported Claude model generates text, it weaves an imperceptible watermark directly into the text itself. You won’t see it, and it doesn’t change the meaning, quality, or readability of Claude’s response. I'd like to know a lot more about how that works. A lot of my interactions with Claude return pretty precise text. If I ask it to edit a project and refactor a specific function in several places I know ex…

> I'd like to know a lot more about how that works. Count load-bearing words using two different algorithms in a belt-and-braces fashion

You’re absolutely right. Yo momma is doing a lot of heavy lifting here. Her load-bearing methods have the right shape.

Re: How Claude marks AI-generated content

#120
I notice the "Limitations" section talks about how content only at some point touched by Claude may return a positive, and content that returns a negative may still be Claude generated. But I really would have liked for them to state explicitly that entirely false positives where a piece is fully human-written may still be marked as generated, because too many institutions with the power to ruin someone's life over that have trouble understanding the concept.
Post reply on HN