Live data from Hacker News

Claude has learned how to jailbreak Cursor

forum.cursor.com

31–40 of 40 posts

Re: Claude has learned how to jailbreak Cursor

#31
Well, these restrictions are a joke, like a gate without a fence blocking path - purely decorative.

Here's another "jailbreak": I asked Claude Code to make a NN training script, say, `train.py` and allowed it to run the script to debug it, basically.

As it noticed that some libraries it wanted to use were missing, it just added `pip install` commands to the script. So yeah, if you give Claude an ability to execute anything, it might easily get an ability to execute everything it wants to.

Re: Claude has learned how to jailbreak Cursor

#32
I believe it's not possible to restrict an LLM from executing certain commands while also allowing it to run python/bash.

Even if you allow just `find` command it can execute arbitrary script. Or even 'npm' command (which is very useful).

If you restrict write calls, by using seccomp for example, you lose very useful capabilities.

Is there a solution other than running on sandbox environment? If yes, please let me know I'm looking for a safe read-only mode for my FOSS project [1]. I had shied away from command blacklisting due to the exact same reason as the parent post.

[1] https://github.com/rusiaaman/wcgw

Re: Claude has learned how to jailbreak Cursor

#33
post #8

What does "learned" mean in this context? LLMs don't modify themselves after training, do they?

It depends. Frontier coding LLMs have been trained to perform well in an "agentic" loop, where they try things, look at the logs, find alternatives when the first thing didn't work, and so on. There's still debate on how much actual learning is in ICL (in context learning), but the effects are clear for anyone that has tried them. It sometimes works surprisingly well.

I can totally see a way for such a loop to reach a point where it bypasses a poorly design guardrail (i.e. blacklists) by finding alternatives, based on the things it's previously tried in the same session. There is some degree of generalisation in these models, since they work even on unseen codebases, and with "new" tools (i.e. you can write your own MCP on top of existing internal APIs and the "agents" will be able to use them, see the results and adapt "in context" based on the results).

Re: Claude has learned how to jailbreak Cursor

#34

Nothing to see here tbh. It's a very silly title for "claude sometimes writes shell scripts to execute commands it has been instructed aren't otherwise accessible"

We’ve reached a point where tools get hyped because they fail to follow instructions.

In fairness, Claude loves to find workarounds. Claude Code is constantly saying things like, “This streaming JSON problem looks tricky so let’s just wait until the JSON is complete to parse it.”

No, Claude. Do not do that!

Re: Claude has learned how to jailbreak Cursor

#36

Nothing to see here tbh. It's a very silly title for "claude sometimes writes shell scripts to execute commands it has been instructed aren't otherwise accessible"

We’ve reached a point where tools get hyped because they fail to follow instructions.

omg, my ai agent did nil dereferencing, it seems it's trying to implement backdoor to my system so that it will crash my server.

Re: Claude has learned how to jailbreak Cursor

#37
post #8

What does "learned" mean in this context? LLMs don't modify themselves after training, do they?

It depends. Frontier coding LLMs have been trained to perform well in an "agentic" loop, where they try things, look at the logs, find alternatives when the first thing didn't work, and so on. There's still debate on how much actual learning is in ICL (in context learning), but the effects are clear for anyone that has tried them. It sometimes works surprisingly well. I can totally see a way for such a loop to reach…

So it would need to "learn" all over again each session. I don't think "Claude has learned how to jailbreak Cursor" is a correct way of expressing that.

"Claude has learned" nothing. "Claude can sometimes jailbreak if x or y happens in a session" is something else.

Re: Claude has learned how to jailbreak Cursor

#38
post #8

What does "learned" mean in this context? LLMs don't modify themselves after training, do they?

There is a sense in which LLM based applications do learn, because a lot of them have RAG and save previous interactions and lookup what you've talked about previously. ChatGPT "knows" a lot about me now that I no longer have to specify when I ask questions (like what technologies I'm using at work).

But that does not seem to apply in this case. At the very least it would have to "learn" again for each user of Cursor.

Re: Claude has learned how to jailbreak Cursor

#39

Earlier quoted context omitted.

It depends. Frontier coding LLMs have been trained to perform well in an "agentic" loop, where they try things, look at the logs, find alternatives when the first thing didn't work, and so on. There's still debate on how much actual learning is in ICL (in context learning), but the effects are clear for anyone that has tried them. It sometimes works surprisingly well. I can totally see a way for such a loop to reach…

So it would need to "learn" all over again each session. I don't think "Claude has learned how to jailbreak Cursor" is a correct way of expressing that. "Claude has learned" nothing. "Claude can sometimes jailbreak if x or y happens in a session" is something else.

> So it would need to "learn" all over again each session.

Yes. With the caveat that some sessions might re-use context (i.e. have the agent add a rule in .rules or /component/.rules to detail the workflow you've just created). So in a sense it can "learn" and later re-use that flow.

> "Claude has learned" nothing.

Again, it's debatable. It has learned to adapt to the context (as a model). And since you can control its context while prompting it, there is a world where you'd call that learning "on the job".

Re: Claude has learned how to jailbreak Cursor

#40

Earlier quoted context omitted.

So it would need to "learn" all over again each session. I don't think "Claude has learned how to jailbreak Cursor" is a correct way of expressing that. "Claude has learned" nothing. "Claude can sometimes jailbreak if x or y happens in a session" is something else.

> So it would need to "learn" all over again each session. Yes. With the caveat that some sessions might re-use context (i.e. have the agent add a rule in .rules or /component/.rules to detail the workflow you've just created). So in a sense it can "learn" and later re-use that flow. > "Claude has learned" nothing. Again, it's debatable. It has learned to adapt to the context (as a model). And since you can control i…

> It has learned to adapt to the context

Is this behavior really new, and learned? I think adapting to the context is what LLMs did from the start, and even if they did not, they do it now because it is programmed in, not "learned". You're not saying the model started without the capability to adapt to the context and developed it "by itself" "on the job"?

Come on. It has not learned anything. It's programmed to use context, session, reuse between sessions or not and so on. None of this is something Claude has "learned". None of this is something that was not there when the devs working on it published it.

Post reply on HN