Live data from Hacker News

How Claude marks AI-generated content

support.claude.com

391–400 of 446 posts

Re: How Claude marks AI-generated content

#391
post #17

Earlier quoted context omitted.

I had a similar thought but I assumed they leaned in because it improved performance on coding or something like that

I suspect it's because of alignment concerns. The more deeply they can integrate their principles, the harder it'll be to misuse. Or at least that's the idea.

From what I heard several AI companies have intentionally been making the personality more cringe so people stop making it their girlfriend.

Re: How Claude marks AI-generated content

#392

Earlier quoted context omitted.

Prompt engineer is a requirement within every serious job now, not a job in itself

I've never known any firefighters to prompt engineer a blaze, but perhaps you don't consider that a "serious" job.

Hey ChatGPT, what side should I make the incision on?

Re: How Claude marks AI-generated content

#393
post #21

An interesting factor of this is competition. If Claude was the only model family they could ship a change like this and users who want to cheat (or don't like watermarks for other reasons) would just have to put up with it. In a world with many different competing models, the risk of losing customers to other providers over this is much more real. Maybe they've looked at the numbers and the portion of people who cle…

Presumably the expected cost of doing it is less than the cost of getting fined by the EU for not doing it.

Re: How Claude marks AI-generated content

#395

I notice the "Limitations" section talks about how content only at some point touched by Claude may return a positive, and content that returns a negative may still be Claude generated. But I really would have liked for them to state explicitly that entirely false positives where a piece is fully human-written may still be marked as generated, because too many institutions with the power to ruin someone's life over t…

> But I really would have liked for them to state explicitly that entirely false positives where a piece is fully human-written may still be marked as generated, because too many institutions with the power to ruin someone's life over that have trouble understanding the concept. This is marketing material aimed, in part, at encouraging the usage you are concerned about, which is why they do not highlight that problem…

But I thought Anthropic was an altruistic organization devoted to the betterment of humanity…

Re: How Claude marks AI-generated content

#396
post #31

Earlier quoted context omitted.

The moment Google announced SynthID, the first domain I bought was deSynthID.com Several open-source projects have already proven SynthID to be ineffective.

There are many free lock picking tutorials, but yet locks are still effective.

this gives strong "you wouldn't steal a car" vibes

Re: How Claude marks AI-generated content

#397
post #215

Earlier quoted context omitted.

I'm curious about your thoughts on pangram. I only really see posts on Reddit claiming it falsely labels their content as ai generated but nobody will actually post examples of "textbook from twenty years ago" or upload screenshots of a journal (also those posts usually feel deeply ai generated without an ai detector) Do you think this is an impossible task and we shouldn't try to solve it? Or do you think it's doabl…

Language distribution shifts. Eventually people will start adopting the distribution used by LLMs, making classification harder. Also, this doesn't even consider the case where people use LLMs to translate their original works. Or people that use it for spelling/grammar checks. Personally, I believe these checkers do more harm than good. Any false positive can ruin someones life.

> Eventually people will start adopting the distribution used by LLMs, making classification harder.

I recently heard someone say "that's genuinely the exact solution I was looking for" and had to do a double take.

Re: How Claude marks AI-generated content

#399

Earlier quoted context omitted.

> But I really would have liked for them to state explicitly that entirely false positives where a piece is fully human-written may still be marked as generated, because too many institutions with the power to ruin someone's life over that have trouble understanding the concept. This is marketing material aimed, in part, at encouraging the usage you are concerned about, which is why they do not highlight that problem…

But I thought Anthropic was an altruistic organization devoted to the betterment of humanity…

It would appear that their altruism isn't very effective.

Re: How Claude marks AI-generated content

#400
post #369

Earlier quoted context omitted.

It’s worse than that, false positives are possible but someone generating text should be able to get ai to change some words and formatting to break the watermarking, then ai detectors can tell them how well they did. I don’t know what the answer but I absolutely know it isn’t this.

I think we need a chain of custody system for content, but that would require browsers, software, websites, operating systems, phones, camera manufacturers, etc to all get on board. But each intermediary or source (optionally) cryptographicaly signs a piece of content that it either generates, edits, or passes along, and the end result at a destination, is that content is either 'trusted' if its cryptographic chain i…

This is a meme video but I think it hits the nail on the head.

TLDR: AI will force global online digital ID for everyone that uses the internet for the exact reason you mentioned. And that would forever change free speech forever allowing the powers that be to put the genie "back in the bottle" so to speak.

link: Raiden Warned About AI Censorship - https://youtu.be/-gGLvg0n-uY

Post reply on HN