Value judgment aside: I am a bit surprised at how sloppily they did this. I think they could've achieved the same effect while decreasing the odds of detection via reverse engineering. (This field is known as "underhanded code", coined by the Underhanded C contest: https://www.underhanded-c.org . It's a little-known "art"; little-known for probably self-explanatory reasons. There are much cleverer ways of achieving o…
It's also possible that there are more in-depth detection methods and that this was just a cheap and easy first step that hasn't been removed because it catches a lot of less sophisticated bad actors. It's unlikely that this will stop a big AI lab from distilling their model if they're really determined, but A) it may be enough to stop a bunch of fly-by-night token resellers looking to make a quick buck and B) you ne…
Claude Code is steganographically marking requests
251–260 of 817 posts
Re: Claude Code is steganographically marking requests
#252There are some commentors in this thread downplaying the severity of a service provider being less than transparent about exactly what their shipped tooling does on customer's machines. That the provider's business needs necessitate the this behaviour doesn't justify their lack of honest disclosure. That honest disclosure would render the solution to their problem useless isn't my problem. If anything, that they thou…
So any covert bullshittery hits hard.
Re: Claude Code is steganographically marking requests
#253Re: Claude Code is steganographically marking requests
#254Earlier quoted context omitted.
That would actually be an interesting thing to read about
Years ago, EVE corps swapped Unicode lookalike characters in patterned ways, inserted patterned zero width space characters, and put very slightly color shifted background watermarks into forum posts to detect leaks.
For "MMO geopolitics fingerprinting", you can in theory do the entire thing mostly or entirely from the server, with the client not actually ever receiving any underhanded code per se. Such as sending dynamic stylesheets that vary in a pretty plausibly deniable way that can be secretly extracted from screenshots. Same for the character swap stuff. A very good analyst could still potentially detect it, but it's much harder.
With this, there's the smoking gun of the semi-deobfuscated underhanded code in the client. It will always have to exist in some form, but you can write it in a way where it not just looks like regular code but actually has a believable purpose and behavior which could plausibly be normal and benign for implementation of a feature or telemetry or whatever. They did not really do it in a sufficiently "cleverly psyop-y" way, so to speak.
Re: Claude Code is steganographically marking requests
#255- filtering out people from the wrong side of "all humanity", years before it was demanded by the government
- downgrading their models in arbitrary ways (later saying "sorry but not really")
- actively sabotaging the replies, as in covertly modifying them to feed the users incorrect results
What's next to expect from Anthropic? Malware to brick your machine if they don't like you? Extending this to more people they don't like? I think I already can see how Dario's Amodei utopian visions of the future of "all humanity" are going to unfold.
Re: Claude Code is steganographically marking requests
#256If there weren't already enough tells that something is AI-generated, I guess you could add this to the list.
Re: Claude Code is steganographically marking requests
#257I’m pretty sure every lab, including Anthropic, is doing distillation right now.
Re: Claude Code is steganographically marking requests
#258Re: Claude Code is steganographically marking requests
#259Re: Claude Code is steganographically marking requests
#260Headline is, frankly, awful. This isn't the AI secretly doing stuff and hiding it. This is the very human Anthropic engineers trying to detect Chinese scraping via some frankly hamfisted and unimaginative URL trickery.