Live data from Hacker News

Claude Code is steganographically marking requests

thereallo.dev

791–800 of 817 posts

Re: Claude Code is steganographically marking requests

#791

Earlier quoted context omitted.

This sounds similar to what people were saying regarding Microsoft when the shady tricks of consumer Windows 10 versions were revealed. …And then Windows 11 became even worse.

Surely we can do better than vague gesturing on HN. For one, stake out a concrete position so people have something of substance to respond to.

You replied to me in the first place… I didn’t seek out one of your comments to reply to?

This ask doesn’t make sense.

Re: Claude Code is steganographically marking requests

#792

What do we think are the chances they trained their models to behave worse or even malicious if those special apostrophes are present in the system prompt?

Degraded performance for resellers and model distillers? Probably, and I don’t think that’s unreasonable on their part. Malicious? I really doubt it.

It'd honestly be pretty alarming. They just took the entirety of the internet to become rich, and now that they have something to take, they intentionally worsen the whole thing (by investing additional training data) to ensure resellers/ distillers get worse responses.

I'm obviously not taking the side of the resellers here - if their T&C don't allow reselling, then by all means restrict their access. But intentionally training the model to give worse responses when it sees a specific date format, would be pretty messed up.

Re: Claude Code is steganographically marking requests

#793

Value judgment aside: I am a bit surprised at how sloppily they did this. I think they could've achieved the same effect while decreasing the odds of detection via reverse engineering. (This field is known as "underhanded code", coined by the Underhanded C contest: https://www.underhanded-c.org . It's a little-known "art"; little-known for probably self-explanatory reasons. There are much cleverer ways of achieving o…

[dead]

Re: Claude Code is steganographically marking requests

#794
post #431

Earlier quoted context omitted.

> Why do you think it's better that a country that turns its citizens into a pulp for criticizing the government, and censors most media to control its citizens' thoughts, reach SI before one that is democratically elected and in which you can generally criticize the government? Which country are you referring to? As an outsider who is neither American or Chinese, day by day it seems like the US is inching towards th…

You cannot be serious if you think this. Please do some actual research on how China treats dissident citizens. They ship their human rights lawyers off to concentration camps, force them to eat their own feces, then rape and/or murder them. 57k citizens disappeared since 2013, at least 5k more every year. I don't get disappeared for criticizing the US government online. The US government doesn't censor most media I…

That's why I said "inching towards", not that it already is.

See how people are being treated in ICE detention camps. See how protests are getting violent with law enforcement officers donning military gear. Your mass media is built for consumption and conformity, not reality. The fact that news channels can say whatever lies they want on commentary programs, washing their hands without any reprisal is telling. You don't need to censor the media if the media is already self-censoring and self-serving due to the fact that it is owned by oligarchs.

What outsiders see is history repeating itself.

And I'm definitely concerned that the US could revoke/not renew my visitor visa based on online posts like this one.

I've been to Beijing International Airport on a layover, and at least the Chinese are on your face about being surveilled. Giant cameras everywhere, to access wifi you had to scan your passport, the Great Firewall is a reality etc.

Meanwhile lots of my personal data is hosted in the US via Google and Apple and I have absolutely no doubt that if some American authority decided to trawl around my files they absolutely could if they follow the (American) due process.

Re: Claude Code is steganographically marking requests

#795
post #249

Earlier quoted context omitted.

You have an odd definition of "blew up in their faces". What, do you somehow think your average Claude Code user on HN is going to think "Oh wow, I'm sure I'll get a much better experience if instead of going to the standard Anthropic Claude API endpoint I go through xiaohongshu.com."

For personal projects with no data sensitivities, I use Claude Code with DeepSeek v4 Pro a lot. I'm probably going to switch to OpenCode or pi.dev after this. I was already a little annoyed at using a closed source harness, but it matched what I used at work. Nowadays, I'm mostly using Codex at work so no reason not to switch anymore.

wait, you can use other models with CC harness?

Re: Claude Code is steganographically marking requests

#796

Earlier quoted context omitted.

> I understand that the model name might appear in prompts for distillation (I guess? "You are RipOffModelv2, learn from these responses from Claude") This does not make sense. You wouldn't send such a prompt to the Claude model. And when you're sending the prompt (anywhere) you don't have the response yet. This is not how distillation works.

Right, sorry, I'm trying to catch up (in general) here, and am working through assumptions to get my bearings. What you say makes sense, but further adds to my confusion as to why those model names would appear in input sent to Claude at all, then. EDIT: I guess it might be because someone might point Claude at a compatible API, with its model in the URL, which is of interest to them.

Yes I think that's the theory. I also noticed something today - when a stakeholder asked if Claude had an affiliate program, my instant response was "no" (because why would they need to deal with that) but a quick search surfaced that there are websites that actively present themselves as "Claude" official sellers/resellers, and simply act as a middleman proxy for the Anthropic/bedrock API.

It's unclear to me if they do this to arbitrage on API costs or to steal data, but it makes sense that Anthropic would be interested in detecting when traffic is being proxied to them.

Re: Claude Code is steganographically marking requests

#797
post #249

Earlier quoted context omitted.

For personal projects with no data sensitivities, I use Claude Code with DeepSeek v4 Pro a lot. I'm probably going to switch to OpenCode or pi.dev after this. I was already a little annoyed at using a closed source harness, but it matched what I used at work. Nowadays, I'm mostly using Codex at work so no reason not to switch anymore.

wait, you can use other models with CC harness?

Yes, as long as they support the Anthropic API standard. You might have luck converting between standards using litellm's proxy. The system prompt and tool choice are tuned for Anthropic's models. For example, this setup will use Deepseek V4 Pro with their first-party (subsidized) API. Things like the Read tool on images won't work, but mostly this works well. Your mileage may vary.

    #!/bin/sh
    export ANTHROPIC_BASE_URL=https://api.deepseek.com/anthropic
    export ANTHROPIC_AUTH_TOKEN=sk-secret
    export ANTHROPIC_MODEL=${ANTHROPIC_MODEL:-deepseek-v4-pro}
    export CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1
    exec claude $@

Re: Claude Code is steganographically marking requests

#798
post #735

Earlier quoted context omitted.

> gaslight users When? > sabatoge their projects Again, when? > attempt a regulatory capture Yet again, when? > ship malware This definitely never happened.

Based only on the third quote (you're literally in the thread discussing second iteration of it), and your username, you can't possibly be acting in good faith here, so I'm not going to waste my time providing references to the events that were widely discussed even here on HN in the past 6 months.

> you're literally in the thread discussing second iteration of it

Did you read the thread before posting this? Half of the thread is about how two bits of information used specifically for abuse prevention are objectively not malware - not that it's necessary to point that out to someone with a functioning brain.

> and your username

Profiling - tell-tale sign of someone too stupid to engage in debate and/or who has zero valid points to contribute.

> I'm not going to waste

Yeah, that fools people below the age of...six? Literally everyone else knows that's just how children desperately try to cover for a lack of ability to provide evidence. "Oh, yeah, my dad owns Epic Games, but because you don't believe me, I'm not going to let you meet him!"

I love triggering the OpenAI shills on HN. It's so trivially easy.

Thank you for creating an indelible record for future HN readers about how every single one of your points is completely unfounded and you're merely a propaganda account. Please continue to respond if you just want to provide further evidence of that.

Re: Claude Code is steganographically marking requests

#799
Even if they really want to detect Chinese, I believe training a classifier to detect from prompts would be very easy for Anthropic (they clearly do not care that much about false positives anyway): Chinese, or even Chinglish, very obvious.

But they would rather plant a Trojan on the user's computer.

Re: Claude Code is steganographically marking requests

#800
post #785

Earlier quoted context omitted.

I didn't make an exceptionalist argument, but if any country's behavior and values can be measured compared to others, you will always be able to make some kind of decision about where those fall in terms of goodness or badness. Do you not believe in good and bad?

I believe in good and bad. I don't believe in US good, non-US bad. I also don't believe the same about my religion or my political party, for that matter. How you measure depends on weights you assign (cultural system of values) and what information you use (media bias). You can rank in the extremes (e.g. North Korea as worse than Belgium), since they come out that way by almost any set of information and values. Com…

Most countries have some kind of story they tell about themselves. New Zealand doesn't have any illusions that it is a superpower. It doesn't have the resources, the talent pool or any of that to even begin to dream of it. That is natural. If New Zealand was powerful, then it would find itself in a position of greater responsibility.

You cannot compare countries that barely have the option of ambition with something like the US and even begin to imagine that it is meaningful. You have the UK, France, Germany, Japan, Russia, China and the US. You can rewind history to name other civilizations.

These kinds of countries are the only ones that matter, because they're the ones that have to answer about what people were thinking when they chose to make use of their power in a way that is relevant on a larger scale. They have to answer about what the reasons were when things went wrong and whether they agree things went wrong. If people are even allowed to talk about it.

It's very cheap to label anything as propaganda without taking the time to appreciate whether it has any merit in terms of the overall behavior of a country or its people. You can always find counter-examples, but how influential are they in the larger picture?

Post reply on HN