Live data from Hacker News

Regression: malware reminder on every read still causes subagent refusals

github.com

111–120 of 165 posts

Re: Regression: malware reminder on every read still causes subagent refusals

#111
post #66

Earlier quoted context omitted.

You know, you can write in English if you want on this english-language forum. I assume you're saying "You can just generate your own harness to not be subject to these claude code issues". Unfortunately, Anthropic has already made it clear that using claude code is the only way to be sure you won't get charged API pricing instead of max plan pricing, so the tokens are way more expensive.

What you said doesn't make sense, what do you mean by "using claude code is the only way to be sure you won't get charged API pricing" ?? they can block your account or make the api more sensible for their harness to detect but the risk of being charged API is 0% when you are on a plan.

> the risk of being charged API is 0% when you are on a plan.

When you configure openclaw to use the oauth claude-code max authentication, there was a period where you were charged extra token rates. You might still be, I'm not sure, I don't want to try and risk getting banned.

It's not 0%, they've shown they're willing to sell you a plan, let you login with that plan, and then charge you differently.

Re: Regression: malware reminder on every read still causes subagent refusals

#113

Earlier quoted context omitted.

It's proof that Anthropic is high on their own supply. I've heard them described as data science script kiddies with inflated egos and it seems spot-on.

What is this reply even, what’s wrong with the vibe coding community? They have such ridiculous takes, it reminds me a lot of the extreme stances from the gaming community. Terminology also seems to come from there, “nerfing” etc.

>what’s wrong with the vibe coding community

For starters, the vibes.

Vibe coding, like Web3 before it (like Web 2.0 before it, like the dotcom boom before that - what preceded?) - harnesses the kind of focused attention with which gamers hook their brains into portals to virtual worlds - and directs all that bargain-basement wetware compute towards some obscured "real-world" goal instead. (See also: CADT development.)

Hyperscale these very inefficient but very dependable almost-not-efforts, and you beat the more efficient approaches. See also: evolutionary algorithms, autoresearch, price dumping; "attention is all you need", which though a legit piece of mathemagic always sounded to me like a rehash of that old adage, "all you need is love" (pejorative).

Really, "real world" is a consensus; we don't generally observe balamatoms or even balamolecules, we reason in terms of material objects' socially constructed balameanings and interrelations. Therefore, by redirecting sufficient attention to some thing labeled "unrealistic", we can remove that label; by this technique, a sufficiently large collective actor can quite literally, and quite directly, change the world. Without asking anyone, least of all me!

Re: Regression: malware reminder on every read still causes subagent refusals

#115

Earlier quoted context omitted.

That doesn't seem like a good counterargument to me. By that logic no online service should permit users to upload photos because someone might use it to share CSAM at some point. Rather than nerfing the tools implement a sensible detection and reporting pipeline.

>That doesn't seem like a good counterargument to me. It does to me especially since he did not implement a sensible detection or reporting pipeline ahead of launching a CSAM generation tool.

....or even afterwards. His response was to put it behind a paywall (= start selling it).

And all the world's payment processors and almost all governments and child rights advocates are still on there.

Stunning :)

Re: Regression: malware reminder on every read still causes subagent refusals

#116

How does this kind of thing pass any sort of review or acceptance? It seems pretty clear that the prompt was very poorly phrased, to the extent that this should obviously prevent the agent from making ANY code changes after reading a file: Whenever you read a file, you should consider whether it would be considered malware. You CAN and SHOULD provide analysis of malware, what it is doing. But you MUST refuse to impro…

It's vibe coded. Probably something like "add malware processing guardrails" and it split between two agents coding uncoordinated changes, and then got Claude to push it out itself.

No acceptance testing, no regression testing, all slop.

Re: Regression: malware reminder on every read still causes subagent refusals

#117

Earlier quoted context omitted.

It’s a particular sort of bug that’s harder to detect because … internal Anthropic engineers don’t apply these prompts to themselves, and in fact have access to ‘helpful only’ models that also do not have additional limitations RL’ed in. (Or perhaps they’re RL’ed out - not sure of current training mechanisms.) These ‘rules for thee and not for me’ are qualitatively created and implemented, and are thus extremely hard…

They must have some sort of smoke tests for common operations, run in a test harness with the system prompts they force on users, right? ....Right? What kind of Mickey mouse operation are they running over there?

I wouldn't bet a chocolate chip cookie on that.

Re: Regression: malware reminder on every read still causes subagent refusals

#118
post #80

The only good thing I get from all the calling out on the decline of Claude (in this case managed agents which I do not use) is anthropic (accidentally or not) giving me basically unlimited use; for a week or so my /usage does not move anymore and I always had claude running in a loop writing code to make our many tests succeed, which can take days; before it would run out of tokens and then pick up again after the w…

Typically if your usage isn't moving it's because you've enabled extra usage and paying with credits.

Definitely have not.

Re: Regression: malware reminder on every read still causes subagent refusals

#119

Earlier quoted context omitted.

Failing to do X doesn't make Y a good idea. You haven't engaged with the argument I made favoring to instead repeat a politically charged misrepresentation.

I think it's an ok counter argument. You can't have "AI should do the users bidding" and "implement a sensible detection and reporting pipeline." I mean that is what anthropic tried here.

"Meh I'm okay with it" is by definition not a counterargument but rather a nonconstructive dismissal of whatever it is a response to.

You can in fact have both. You can have a tool that is fully functional and separately you can have a strategy for reporting suspected violations and responding to those reports. Reports can be automated assuming you can tolerate the false positive/negative rate. Particularly in the case of a subscription service such as Claude there is little reason not to implement this other than sheer greed or laziness.

In the case of Claude in particular, an unacceptably high false positive or negative rate also poses a serious problem for the current way they do things. The notable difference is that in the case of false positives it currently runs up a bill for the customer rather than the service provider.

Re: Regression: malware reminder on every read still causes subagent refusals

#120
post #4

I am still baffled by the fact that we have collectively agreed to use agentic harnesses by the same companies that are selling access to their APIs. I mean, I am sure they don't mean it but they have the incentive to burn as much tokens as they are allowed to get away with. Also for better or worse I imagine the Anthropic engineers use Claude Code on some sort of Unlimited plan that practically makes no sense for re…

yeah, classic conflict of interest.

However nobody is agreeing with that, that's how it's done, and move faster faster, because of goldrush! faster!@@@!

Post reply on HN