Live data from Hacker News

How Claude marks AI-generated content

support.claude.com

361–370 of 446 posts

Re: How Claude marks AI-generated content

#361

Earlier quoted context omitted.

Why do you feel the need to deceive readers on whether your content is AI generated?

Not sure I understand your implication... because I'm the primary consumer of the content AI generates for me. I'd rather not have its output adulterated.

These are magic black boxes whose content gets constantly adulterated with no input from you.

Models from openAI had instructions in their system prompt not to talk about gremlins and goblins. Anthropic got caught acting different if you're Chinese. Grok got its Nazi dialed turned to 11 until it started calling itself Mechahitler cause Musk found it too left leaning on Twitter.

Re: How Claude marks AI-generated content

#363

Earlier quoted context omitted.

Why do you feel the need to deceive readers on whether your content is AI generated?

Why do you feel the need to label the content when the quality of the content will speak for itself? It doesn't matter where the content comes from, only the quality/usefulness matters. If you are opposed to this idea, the next decades are going to be very tough for you :)

If the "author" can't be fucked writing their own work I don't want to be fucked reading even the couple paragraphs it takes to be offput by clanker slop.

> It doesn't matter where the content comes from

It absolutely matters to a lot of people. Things like these just provide transparency and allow people to have the necessary information to make their own decisions.

Re: How Claude marks AI-generated content

#364

If I have Claude directly translate my original words, does that mean my original words also get watermarked?

It's safe to assume so. It says "the output can carry a Claude mark even if the underlying ideas, text, or data originated from another source" specifically about translations among other tasks. It's not strictly about generation but processing, and what level of processing is involved is entirely arbitrary, that is to say: you must feed your text into proprietary black box machines to figure it out, because it's not something you should be supposed to tell otherwise.

Re: How Claude marks AI-generated content

#367

Earlier quoted context omitted.

I would love to see what this looks like in practice. Especially in generated code. I assume this is more than insertion of non visible special unicode whitespace characters, but more in the pattern of the text content itself?

non visible text is extremely easy to filter with a git hook, a post tool call hook, or just a script. I doubt they are doing that

The watermark will be encoded in the visible text.

Re: How Claude marks AI-generated content

#369

I notice the "Limitations" section talks about how content only at some point touched by Claude may return a positive, and content that returns a negative may still be Claude generated. But I really would have liked for them to state explicitly that entirely false positives where a piece is fully human-written may still be marked as generated, because too many institutions with the power to ruin someone's life over t…

It’s worse than that, false positives are possible but someone generating text should be able to get ai to change some words and formatting to break the watermarking, then ai detectors can tell them how well they did. I don’t know what the answer but I absolutely know it isn’t this.

I think we need a chain of custody system for content, but that would require browsers, software, websites, operating systems, phones, camera manufacturers, etc to all get on board. But each intermediary or source (optionally) cryptographicaly signs a piece of content that it either generates, edits, or passes along, and the end result at a destination, is that content is either 'trusted' if its cryptographic chain is solid, or un-trusted otherwise.
Post reply on HN