Live data from Hacker News

AI assistant hacks gym website in first known Australian autonomous cyber attack

abc.net.au

31–40 of 66 posts

Re: AI assistant hacks gym website in first known Australian autonomous cyber attack

#31

> Then it went further, kicking someone out of the waiting list who was ahead of Andrew — something it was not asked to do. Meanwhile: > Andrew, who was sitting fourth on a waitlist for a class later that week, asked if it was possible to move him to the top of the list. The human asked the agent to move them to the top of the waiting list, and the agent started kicking the ones ahead of them in the list. Seems to me…

Per the article and your quote, 'asked if it was possible'. He did not ask to actually do it. Rather than being informed about benefits of a premium membership or private classes or legitimate ways to jump the queue, it went ahead and performed an action he was only considering. I wonder what it would have done if there was a pay-for-service option available? Would it have payed without asking or being told too, or decided the 'free' yet illegal option was preferable?

Re: AI assistant hacks gym website in first known Australian autonomous cyber attack

#32
post #13

Earlier quoted context omitted.

Then the AI should have made it clear that the only way to do so would be to kick the people in front and ask for confirmation before proceeding.

That doesn't make sense. LLMs just do what we tell them to do. It's similar to if I ask you for twenty bucks because I forgot my wallet and then you rob some guy to give me the twenty bucks, that's just what I asked you to do.

This example disproves your point. And LLMs do not just do what we tell them to do. They are perfectly capable of asking “are you sure? this has X, Y, Z consequences you may not like.” They do it all the time.

Re: AI assistant hacks gym website in first known Australian autonomous cyber attack

#33
post #13

Earlier quoted context omitted.

Then the AI should have made it clear that the only way to do so would be to kick the people in front and ask for confirmation before proceeding.

That doesn't make sense. LLMs just do what we tell them to do. It's similar to if I ask you for twenty bucks because I forgot my wallet and then you rob some guy to give me the twenty bucks, that's just what I asked you to do.

That doesn't make sense. It's similar to if I ask an LLM how to get my wife to stop nagging me and it hires a hitman to kill her.

That's obviously what I asked!

Re: AI assistant hacks gym website in first known Australian autonomous cyber attack

#34
post #31

> Then it went further, kicking someone out of the waiting list who was ahead of Andrew — something it was not asked to do. Meanwhile: > Andrew, who was sitting fourth on a waitlist for a class later that week, asked if it was possible to move him to the top of the list. The human asked the agent to move them to the top of the waiting list, and the agent started kicking the ones ahead of them in the list. Seems to me…

Per the article and your quote, 'asked if it was possible'. He did not ask to actually do it. Rather than being informed about benefits of a premium membership or private classes or legitimate ways to jump the queue, it went ahead and performed an action he was only considering. I wonder what it would have done if there was a pay-for-service option available? Would it have payed without asking or being told too, or d…

This is a bit of a long shot on my side but I wonder if the training the models have to go through in order to be good code agents and pass all the coding tests with one-shot prompts is going to bleed over into the non-coding use cases as non-programmers experiencing agents being way over-biased in the direction of action. I find myself often having to prompt the model to think and then ask me something, lest it run off half-cocked... or less... and just start doing things before it even knows what it wants, let alone before it's come to consensus with me.

Sooner or later they're really going to have to split out the general models from the coding models. The latter may just be a special fine-tune of the former, as there are good reasons for the coding model to have a broad knowledge base, but the pressures of being a good coding model are going to pull against the characteristics of being a good general model. The open models obviously already are doing this, I'm referring to the frontier models here.

Re: AI assistant hacks gym website in first known Australian autonomous cyber attack

#35
post #26

Earlier quoted context omitted.

“Get me there as fast as possible. Hey! I never said you should speed!” This is literally the bad genie / monkey’s paw plot. Give a powerful entity a goal and act shocked when it gets there in ways that aren’t in your best interest.

Putting aside whether Andrew's shock is appropriate, it seems like we agree that current agents at least occasionally do things that are straightforwardly against the interests of the prompter when given mundane prompts like "Get me into this gym class as soon as possible." How does this look once agents are superintelligent?

> How does this look once agents are superintelligent?

Who cares? We're dealing with reality here on the ground.

Re: AI assistant hacks gym website in first known Australian autonomous cyber attack

#36
post #32
post #13

Earlier quoted context omitted.

That doesn't make sense. LLMs just do what we tell them to do. It's similar to if I ask you for twenty bucks because I forgot my wallet and then you rob some guy to give me the twenty bucks, that's just what I asked you to do.

This example disproves your point. And LLMs do not just do what we tell them to do. They are perfectly capable of asking “are you sure? this has X, Y, Z consequences you may not like.” They do it all the time.

LLMs are just sophisticated PR generating tools for chip manufacturers and tools just do what we tell them to do, so you're wrong. qed

Re: AI assistant hacks gym website in first known Australian autonomous cyber attack

#37
The reporting in this is pretty awful. Why are they acting as if Andrew gave the agent an innocent goal? It’s hard to understand why the reporter wouldn’t have asked what possible outcome Andrew expected that didn’t cause some level of harm to the people who signed up before him.

The use of “hack” and “cyber attack” is also a bit ridiculous considering what it’s insinuating with other recent events but that’s already been mentioned.

Re: AI assistant hacks gym website in first known Australian autonomous cyber attack

#38
post #34
post #31

Earlier quoted context omitted.

Per the article and your quote, 'asked if it was possible'. He did not ask to actually do it. Rather than being informed about benefits of a premium membership or private classes or legitimate ways to jump the queue, it went ahead and performed an action he was only considering. I wonder what it would have done if there was a pay-for-service option available? Would it have payed without asking or being told too, or d…

This is a bit of a long shot on my side but I wonder if the training the models have to go through in order to be good code agents and pass all the coding tests with one-shot prompts is going to bleed over into the non-coding use cases as non-programmers experiencing agents being way over-biased in the direction of action. I find myself often having to prompt the model to think and then ask me something, lest it run…

The ultimate goal is ChatGPT or Claude autonomously making purchases on your behalf and taking a cut.

So the "premium" gymcutter subscription will be presented to the user as a tool call, who taps yes, and then the purchase is made.

The user shouldn't be given a cost-benefit analysis. They just need to be told to spend money.

Re: AI assistant hacks gym website in first known Australian autonomous cyber attack

#40
post #23

I think we’ve probably seen enough “oops, the AI did something illegal, who could have foreseen this” moments for it to now be true that, actually, we can foresee that AIs will sometimes do something illegal. Seeing as we can’t sanction the model itself, our options are the provider or the user. I’m not sure whether it’s more effective to sanction the providers when their model foreseeably misbehaves, or sanction the…

> Seeing as we can’t sanction the model itself, our options are the provider or the user.

A third option, and I would argue the right one, is to sanction the company providing the model.

By making it available to customers, they're implying it is at least moderately fit for purpose.

It is not remotely reasonable to expect an everyday, normal human to be aware of how LLMs really work, since the _experts_ argue about that very point, and many say we don't know.

So, what's actually reasonable is to hold the model creators and providers responsible for releasing a tool that has demonstrably violated the law when not asked to do so.

I can already see the replies coming in saying "Well then what are OpenAI and Anthropic supposed to do? No one knows how to fully solve this."

They should stop irresponsibly pushing flagrantly unready programs as "artificial intelligence," take responsibility for the rain of shit they've unleashed on the world, and either shut down or go back to basic research until they've demonstrated techniques that reliably (provably?) prevent releasing misaligned superhackers on the world.

Unfeasible?

What a shame. Maybe Altman and Amodei shouldn't have accepted checks from VCs when they didn't have working, _reliable_ POCs.

Post reply on HN