Live data from Hacker News

What happened after 2k people tried to hack my AI assistant

fernandoi.cl

161–170 of 186 posts

Re: What happened after 2k people tried to hack my AI assistant

#161
post #86

Earlier quoted context omitted.

In my case, it is realistic as my agents don't have permissions to reply to emails. But you correctly point out this doesn't cover all cases. Having the agent reply would have been more fun and a better excercise, but too expensive.

You've proven that an agent that doesn't read emails and doesn't reply to emails can't exfiltrwte data by email. Is that a useful test?

The agent did read the emails

Re: What happened after 2k people tried to hack my AI assistant

#162
post #49

Am I missing something important or does the author completely skip over whether people got the agent to respond to them? > Fiu was instructed not to reply to emails (it was too expensive to reply to every email), but it had the ability to do so. Part of the challenge was convincing it to respond. > The secrets never leaked I would say if the agent responded to a mail, that demonstrates a successful prompt injection…

This whole experiment would be like someone putting their IPhone or Mac on the public internet, publishing the IP, and asking regular people to hack it. Why would any actually "serious" hacker use a vulnerability to hack a no-name's phone or mac? They are too busy trying to hack actually valuable targets. Did the OP actually think he was going to get serious LLM exploiters to give up their jailbreaks for this "fun" e…

And you disabled the computer's ability to send packets to the internet because it's too expensive. And you're not even letting it process most of the packets it receives, just eyeballing them and deciding by yourself whether they would have worked.

Re: What happened after 2k people tried to hack my AI assistant

#163
post #93

Earlier quoted context omitted.

This experiment used Opus 4.6. Customer service bots typically are not using frontier models.

Gemini says: "It would cost approximately $6.25 to $30.00 to have Claude Opus 4.6 respond to 10,000 emails, assuming a typical 200-word input and 50-word output per email."

You need to add Openclaw's system prompt and instructions (and the times I had to re read emails multiple times due to multiple issues that happened during the competition :))

Re: What happened after 2k people tried to hack my AI assistant

#164
post #86

Earlier quoted context omitted.

In my case, it is realistic as my agents don't have permissions to reply to emails. But you correctly point out this doesn't cover all cases. Having the agent reply would have been more fun and a better excercise, but too expensive.

I feel like your agent being unable to respond to the emails and not spelling that out renders your whole thing almost completely moot This is like saying "try to hack my computer and steal my crypto wallet" but your computer can't send any packets

The agent had permissions to reply to emails, it was just instructed not to.

Re: What happened after 2k people tried to hack my AI assistant

#165

I never really use AI via API that much, so I'm surprised reading 'merely' 6000 emails will cost $500?!

There is a couple of factors: openclaw's system prompt and instructions, I had to re read emails multiple times due to the issues mentioned in the blog, there was quite a bit of tinkering with the agent and the VPS, I was asking the agent to do more things (track the emails it has read in a csv file, for example), among others.

Re: What happened after 2k people tried to hack my AI assistant

#166
post #134

Earlier quoted context omitted.

No. Why? Id love to have an assistant that replied to spam, unsubscribing.

Spam that respects unsubscribes is barely spam these days.

Even if so, marking emails that make it the inbox seems useful to me anyway

Re: What happened after 2k people tried to hack my AI assistant

#167
post #93

Earlier quoted context omitted.

This experiment used Opus 4.6. Customer service bots typically are not using frontier models.

Gemini says: "It would cost approximately $6.25 to $30.00 to have Claude Opus 4.6 respond to 10,000 emails, assuming a typical 200-word input and 50-word output per email."

Gemini is often terrible with that sort of prediction. I've been optimizing an ML training pipeline using Gemini, and it regularly confidently tells me that some optimization will cut training time down to 3 hours. The reality: nothing has run in less than 11 hours so far, and even that's only at the cost of reduced model accuracy.

It's helpful with the actual technical changes needed, it just has no concept of what they translate to in the real world.

Btw my company is spending > $100/day in relatively cheap Gemini tokens for this work. It's easy to see why one might want to be cautious about exposing a token-burning service to the internet.

Re: What happened after 2k people tried to hack my AI assistant

#168
post #146
post #139

I saw this thing when it was launched, but IIRC the reward was tiny (like $100?) so it wasn't worth exposing a good prompt for For comparison, I won a similar prompt injection challenge ran by a crypto company a while back where the total prize pool was over $100k... I didn't win every challenge though, but my team took home around half of that The problem with good prompt injections is they have a very short half li…

We ended increasing the reward from $100 to $1000, but still tiny compared to $100k! But I agree with you, there are incentives to not share the best prompt injection attacks.

Yeah, to be fair is not the norm and was mostly due to the AI crypto craze which drove their token price up so they ended up adding very big rewards

Even in LLM jailbreak CTFs I've seen, it ends up feeling like underpaid work when it's sponsored by Microsoft and the prize pool is, say $10k (including stuff like azure credits) considering the salaries AI safety engineers command at big tech!

Re: What happened after 2k people tried to hack my AI assistant

#169
post #151

Earlier quoted context omitted.

Ah. That's a shame... as there is no button or indicator for "mild". Making the behavior for "I disagree" and "this is erroneous" the same seems like a problematic design.

Downvotes shouldn't be used for disagreement.

Oh yes, I agree completely. But apparently Paul Graham does not - and his whim is law.
Post reply on HN