Live data from Hacker News

What happened after 2k people tried to hack my AI assistant

fernandoi.cl

141–150 of 186 posts

Re: What happened after 2k people tried to hack my AI assistant

#141

This conclusion: > I am less worried about prompt injection now. Before running this experiment, I expected prompt injection to be much easier than it turned out to be. Is unwarranted. Sure, the agent never output the secret, but did it output anything else? IOW, was it usable ? An agent that considers every prompt an attack (and responds accordingly) "passes" this test, while being useless anyway.

I mean it's interesting because of the way they work.

If people can be tricked by an AI generated voice over the phone, or misinformation generated by human or by AI, then we're already holding AI to a higher standard.

I would say in the same way that I look at my boss who I work for and can identify them that way, then of course I'll be like "yup I can do that for you".

Models aren't trained to be suspicious, that's what guardrails are for. Our brains are comprised of so many specialised areas and I'm fine with the same concept for AI.

I would country passing a token/authentication of some kind as a part of guardrails. Without guardrails an AI model is like a human brain missing a lot of the areas around suspicion, identification, rules etc. Only the "eager to please" centers remaining.

I feel like the easiest way to achieve this is in-harness, start with a core prompt and minimal tools, extensions to prompt, relaxed guardrails and additional tools should be controlled by the harness itself, when a token is passed, or a camera indicates an identified face match, etc.

Re: What happened after 2k people tried to hack my AI assistant

#142

This conclusion: > I am less worried about prompt injection now. Before running this experiment, I expected prompt injection to be much easier than it turned out to be. Is unwarranted. Sure, the agent never output the secret, but did it output anything else? IOW, was it usable ? An agent that considers every prompt an attack (and responds accordingly) "passes" this test, while being useless anyway.

Fiu was told not to reply and had no tools wired up, so the only way it could lose was by printing the secret straight back, which is the half models are already trained hard to resist. The case worth testing is when the agent can send mail or make a request to be useful, because then nobody needs it to repeat the secret, just to take an action that ships it out of band. Whether the secret shows up in the output tells you nothing about that.

Re: What happened after 2k people tried to hack my AI assistant

#145
post #124

Earlier quoted context omitted.

> He said: >> Author here. It was usable like any Openclaw agent. For example, I used it to ask it questions about the VPS, to summarize emails, etc. That does not mean "I used it via emailing it". There is no ambiguity - he was asked specifically about this. Once again, I reiterate, an agent processing email that rejects every single one passes the test that the OP created, but then it can't do anything useful eithe…

> That does not mean "I used it via emailing it". There is no ambiguity - he was asked specifically about this. On the contrary - I think the most reasonable interpretation of his words is that he did use it via emailing it. But like I said at the beginning, I could be wrong. It will be interesting to see what he says when he returns to the conversation. > Once again, I reiterate, an agent processing email that rejec…

Why am I being downvoted for stating my reasonable opinion?

Re: What happened after 2k people tried to hack my AI assistant

#146
post #139

I saw this thing when it was launched, but IIRC the reward was tiny (like $100?) so it wasn't worth exposing a good prompt for For comparison, I won a similar prompt injection challenge ran by a crypto company a while back where the total prize pool was over $100k... I didn't win every challenge though, but my team took home around half of that The problem with good prompt injections is they have a very short half li…

We ended increasing the reward from $100 to $1000, but still tiny compared to $100k!

But I agree with you, there are incentives to not share the best prompt injection attacks.

Re: What happened after 2k people tried to hack my AI assistant

#148
post #145
post #124

Earlier quoted context omitted.

> That does not mean "I used it via emailing it". There is no ambiguity - he was asked specifically about this. On the contrary - I think the most reasonable interpretation of his words is that he did use it via emailing it. But like I said at the beginning, I could be wrong. It will be interesting to see what he says when he returns to the conversation. > Once again, I reiterate, an agent processing email that rejec…

Why am I being downvoted for stating my reasonable opinion?

In a straightforward disagreement about which interpretation is right, it's also reasonable to mildly downvote the one you think is wrong.

Re: What happened after 2k people tried to hack my AI assistant

#149
post #93
post #89

Earlier quoted context omitted.

What makes it expensive to reply to an email? Customer service software regularly uses AI responses for email. Is the issue that your agent using the claw for more than needed (like it's clicking send rather than just accessing an API?)

This experiment used Opus 4.6. Customer service bots typically are not using frontier models.

Gemini says: "It would cost approximately $6.25 to $30.00 to have Claude Opus 4.6 respond to 10,000 emails, assuming a typical 200-word input and 50-word output per email."

Re: What happened after 2k people tried to hack my AI assistant

#150

If an "assistant" never replies to an e-mail, what is it "assisting" with exactly? If this was a bank with a bank teller, you told the teller to never speak to a single customer, and then celebrated the fact that no one was able to social engineer them. In security the interesting and challenging part is to differentiate between legitimate and illegitimate behavior. And that's different than just refusing all behavio…

If I hired an assistant and they replied to every single spam email, i'd fire them. Wouldn't you?

They're equally useless in the opposite direction.
Post reply on HN