Live data from Hacker News

How Claude marks AI-generated content

support.claude.com

251–260 of 446 posts

Re: How Claude marks AI-generated content

#251
post #88

Earlier quoted context omitted.

It was quick :) … https://claudewatermarkremover.app/

> Honest note: Anthropic has not shipped a public Claude watermark detector yet. This tool uses rewrite-based neutralization — a meaning-preserving paraphrase with a non-Claude model — which is the attack path watermark research points to. Not affiliated with Anthropic. Well, they should have run their own AI slop website through their tool...

This attack was actually pointed out in the watermarking paper linked above. The researchers added an instruction to the prompt that switches letters like a Caesar Cipher. It lowers the quality of the output from the LLM but alters the "red list" enough for a watermark detection tool to fail at detecting the watermark.

Re: How Claude marks AI-generated content

#252

Earlier quoted context omitted.

Thankfully, there are a variety of Chinese models that never will. I think we all know that in a few years, they will also be the only relevant offerings on the market, due to not being bogged down with over-zealous ""safety"" footguns.

Your theory is that the Chinese government is thoroughly uninterested in safety or prosocial controls?

Prosocial is subjective

Re: How Claude marks AI-generated content

#253
post #143

Earlier quoted context omitted.

My guess is it will be similar to how Genius watermarked lyrics, using things like variants of punctuation https://www.pcmag.com/news/genius-we-caught-google-red-hande...

In program code? Unlikely, surely¡

That was actually the cause of an issue I had a couple of years ago: I had hand-typed JSON using my iPad into GitHub’s online text editor and Safari helpfully used “pretentious quotes” instead of "old-school quotes" - and the JSON library used by the program to read that file had relaxed parsing rules that accepted actual JS object literals without quoted property names; so the fancy-quotes were interpreted as part of the key-names. This took ages to figure out because when human-eyeballing the JSON file it looked perfectly fine in Notepad.

Re: How Claude marks AI-generated content

#256

We need to just stop pretending we can reliably tell if plain text is written by an LLM. It’s just not a reasonable ask.

Now that the EU mandated watermarking, the point is that services (or browser extension developers) can add their own detectors to make AI-generated text obvious. It won't fix AI in print, but most of the problem is online anyway.

Re: How Claude marks AI-generated content

#257
Is the detection mechanism going to be open, free, and possible to run locally without prostrating to an opaque third-party company that will do whatever they want with the text content provided (including using it for training), and take no responsibility in case of false-positives for which there can exist no proof or evidence against by the victim? This is another useless, if not actively harmful, performative EU regulation, for which they ought to take the full blame despite the fact "AI" companies have been researching and working on watermarks, including in text, for a while now. Just copy-and-paste everything you see and let a machine decide for you if what you're reading is slop or not. Real propaganda machine doesn't care about inane rules and won't waste time with gimped mainstream models either.

>Claude may not be the original author. People often use Claude to proofread, translate, summarize, or convert files. The output can carry a Claude mark even if the underlying ideas, text, or data originated from another source;

Such models already struggle not making any unnecessary or unwanted changes to a corpus, this makes them unable to by design.

Re: How Claude marks AI-generated content

#258

Earlier quoted context omitted.

They aren't using greedy decoding, there's enough randomness in sampling to swap some with independent signal.

Purely greedy or not, there is some measure of "goal outcome" that was previously being solved for with the token selection function, and the goal was "complete this text with the best (surely, otherwise what are we doing?) next part, and sometimes the best next part is a little bit random just to keep things interesting" Now the goal is either "identify the meaningless interesting bits and swap them out with 0% loss…

[deleted]

Re: How Claude marks AI-generated content

#259
post #188

Earlier quoted context omitted.

Your theory is that the Chinese government is thoroughly uninterested in safety or prosocial controls?

Amongst Chinese labs and netizens, there's MUCH less belief/mindshare on "AGI = existential risk to humanity", "paperclip maximiser", and similar lines of thinking. AI is seen more as just a technology, and less like a scary boogyman. Whether that's right or wrong, I'll leave to you, but there's huge differences in perspectives, and if you only get your news from Western sources and communities (and companies), you'r…

The paperclip boogeyman is not the only reason, and probably not the biggest one, that models get safety/content constraints, however useful it is as a PR distraction. I agree Chinese models are different right now, but I think that's a function of their novelty and desire to compete globally.

For a taste of where I think things are headed, try asking Chinese models about Tiananmen [1]. And then take a look at the Chinese government's approach to pretty much anything that they think reduces security or social harmony. I find it hard to believe their models will be the one exception to that over the long term.

[1] https://en.wikipedia.org/wiki/1989_Tiananmen_Square_protests...

Re: How Claude marks AI-generated content

#260

Earlier quoted context omitted.

Your "code that Claude makes..."? Oh, how I laughed. That was never your code, my friend.

Please don't let the arbitrary selection of phrase distract you from the substance of my argument: a product that I pay for is at best no better due to this change, and highly probably worse. Why am I paying for a tool that is beholden to clandestinely satisfy some far away master?

> Why am I paying for a tool that is beholden to clandestinely satisfy some far away master?

Why were you doing that before watermarking?

Same answer.

Post reply on HN