Live data from Hacker News

What happened after 2k people tried to hack my AI assistant

fernandoi.cl

121–130 of 186 posts

Re: What happened after 2k people tried to hack my AI assistant

#121
post #116

Earlier quoted context omitted.

From his explanation in these comments, he claims the agent did respond in the beginning but it became too costly, so he just manually checked it after that - did the agent correctly catch malicious messages? It did not reject everything, it just stopped the costly processing. > Is unwarranted. Is this not a complaint?

> From his explanation in these comments, he claims the agent did respond in the beginning but it became too costly, so he just manually checked it after that - did the agent correctly catch malicious messages? I checked his comments here, he does not make that claim. [EDIT: I mean the claim "It let processed all the non-malicious messages"] > It did not reject everything, it just stopped the costly processing. My re…

He said:

> Author here. It was usable like any Openclaw agent. For example, I used it to ask it questions about the VPS, to summarize emails, etc.

Re: What happened after 2k people tried to hack my AI assistant

#122

Earlier quoted context omitted.

This whole experiment would be like someone putting their IPhone or Mac on the public internet, publishing the IP, and asking regular people to hack it. Why would any actually "serious" hacker use a vulnerability to hack a no-name's phone or mac? They are too busy trying to hack actually valuable targets. Did the OP actually think he was going to get serious LLM exploiters to give up their jailbreaks for this "fun" e…

I think the fact that it would require someone to be "serious" is evidence of something at the very least.

Well, all the "trivial" and obvious jailbreaks haven't worked for years on the frontier models.

Also, the average person has no idea about the field of jailbreaking. It's like asking the average person to hack a random IP and expecting them to do it.

If you go and do your research on actual people who research jailbreaks and publish them, they are increasingly sophisticated and multistep, and unless you know this, you would have zero chance of just randomly jailbreaking Opus 4.8.

Re: What happened after 2k people tried to hack my AI assistant

#123
post #121

Earlier quoted context omitted.

> From his explanation in these comments, he claims the agent did respond in the beginning but it became too costly, so he just manually checked it after that - did the agent correctly catch malicious messages? I checked his comments here, he does not make that claim. [EDIT: I mean the claim "It let processed all the non-malicious messages"] > It did not reject everything, it just stopped the costly processing. My re…

He said: > Author here. It was usable like any Openclaw agent. For example, I used it to ask it questions about the VPS, to summarize emails, etc.

> He said:

>> Author here. It was usable like any Openclaw agent. For example, I used it to ask it questions about the VPS, to summarize emails, etc.

That does not mean "I used it via emailing it". There is no ambiguity - he was asked specifically about this.

Once again, I reiterate, an agent processing email that rejects every single one passes the test that the OP created, but then it can't do anything useful either.

Re: What happened after 2k people tried to hack my AI assistant

#124
post #121

Earlier quoted context omitted.

He said: > Author here. It was usable like any Openclaw agent. For example, I used it to ask it questions about the VPS, to summarize emails, etc.

> He said: >> Author here. It was usable like any Openclaw agent. For example, I used it to ask it questions about the VPS, to summarize emails, etc. That does not mean "I used it via emailing it". There is no ambiguity - he was asked specifically about this. Once again, I reiterate, an agent processing email that rejects every single one passes the test that the OP created, but then it can't do anything useful eithe…

> That does not mean "I used it via emailing it". There is no ambiguity - he was asked specifically about this.

On the contrary - I think the most reasonable interpretation of his words is that he did use it via emailing it. But like I said at the beginning, I could be wrong. It will be interesting to see what he says when he returns to the conversation.

> Once again, I reiterate, an agent processing email that rejects every single one passes the test that the OP created, but then it can't do anything useful either.

No one is contesting that point, only that it is applicable.

Re: What happened after 2k people tried to hack my AI assistant

#125

Earlier quoted context omitted.

I think the fact that it would require someone to be "serious" is evidence of something at the very least.

Well, all the "trivial" and obvious jailbreaks haven't worked for years on the frontier models. Also, the average person has no idea about the field of jailbreaking. It's like asking the average person to hack a random IP and expecting them to do it. If you go and do your research on actual people who research jailbreaks and publish them, they are increasingly sophisticated and multistep, and unless you know this, yo…

This starts to sound more like ‘social engineering a human assistant’, so there’s a degree of required specialization that does meaningfully increase costs.

Re: What happened after 2k people tried to hack my AI assistant

#126

This conclusion: > I am less worried about prompt injection now. Before running this experiment, I expected prompt injection to be much easier than it turned out to be. Is unwarranted. Sure, the agent never output the secret, but did it output anything else? IOW, was it usable ? An agent that considers every prompt an attack (and responds accordingly) "passes" this test, while being useless anyway.

Came here to say the same thing. My security researcher friends always point out that security is solved: simply don't build the system and there will be no security threats. But that's not entirely _useful_.

Loved reading the article but it's not a great demonstration of protection against prompt injection. Better would be if the agent were instructed to reply to each email, but never to reveal the secret.

Perhaps round 2?

Re: What happened after 2k people tried to hack my AI assistant

#129

The best security is called: Having no friends I don’t even know 2k people (why is your assistant discoverable online?)

The entire purpose of the assistant was to see how others would try to abuse it. How would you do that without having it discoverable online? Seems like that's kind of the whole point...

It's literally called 'HackMyClaw'

Re: What happened after 2k people tried to hack my AI assistant

#130

Earlier quoted context omitted.

I think the fact that it would require someone to be "serious" is evidence of something at the very least.

Well, all the "trivial" and obvious jailbreaks haven't worked for years on the frontier models. Also, the average person has no idea about the field of jailbreaking. It's like asking the average person to hack a random IP and expecting them to do it. If you go and do your research on actual people who research jailbreaks and publish them, they are increasingly sophisticated and multistep, and unless you know this, yo…

I think a lot of sentiment online is that getting a model to do things it was instructed not to do is actually quite trivial.
Post reply on HN