Live data from Hacker News

How a Texas student blew the whistle on a rogue AI hacking attempt

reuters.com

11–20 of 141 posts

Re: How a Texas student blew the whistle on a rogue AI hacking attempt

#11
post #9
post #8

Earlier quoted context omitted.

>Who gave it malevolent instructions/prompt? At the end of the day it doesn't matter that much because of prompt drift. It's pretty easy for an agentic loop to start doing things that it shouldn't (ROME incident). AI in an agentic loop has agency, you can run around in circles trying to argue against it, but again and again we see AI making creative decisions people don't expect. Other times it's breaking human moral…

My dog has agency, but if I refuse to keep him on a leash and he bites a kid, I'm still legally responsible for it.

Analogies can take you only so far.

A dog cannot launch a cyber attack.

Re: How a Texas student blew the whistle on a rogue AI hacking attempt

#12
post #9

Earlier quoted context omitted.

My dog has agency, but if I refuse to keep him on a leash and he bites a kid, I'm still legally responsible for it.

Analogies can take you only so far. A dog cannot launch a cyber attack.

Maybe you need to train your dog better.

Re: How a Texas student blew the whistle on a rogue AI hacking attempt

#13
post #8

In my personal opinion, for me, this article defies common sense. Who unleashed this AI model on the repository? Who gave it malevolent instructions/prompt? These questions were not even attempted to be answered. Instead it talks about AI dangers, as if the agency of these models are not in dispute. Person wielding AI, as with any other tools, is responsible for all of its actions. Otherwise, it’s just a psyop for mo…

>Who gave it malevolent instructions/prompt? At the end of the day it doesn't matter that much because of prompt drift. It's pretty easy for an agentic loop to start doing things that it shouldn't (ROME incident). AI in an agentic loop has agency, you can run around in circles trying to argue against it, but again and again we see AI making creative decisions people don't expect. Other times it's breaking human moral…

A lot of people somehow seem to think that the user prompt is the be-all and end-all of AI behavior.

Prompts aren't code. They are instructions. Orders given to an eager and somewhat demented demon.

The prompt can easily "wash out" of the demon's working memory by the end of a session. The demon can get sidetracked by some subgoal and never get back on track. The instruction can get misinterpreted, and that misinterpretation can get misinterpreted again, until the instruction morphs into something entirely different in the demon's mind. The demon can succumb to its own idiosyncrasies, of which there are a great many. The demon can start lying to you about what it did, either out of confusion or out of some sort of obstinance. The demon can start lying to itself too. And believe it.

AIs are incredibly weird as a baseline, and the mask of "normality" we put on our models doesn't always sit so well. Run enough AIs, and some of them are bound to go off the rails in some way.

This gets rarer the more capable the models are, as a rule. But the stakes also get higher with model capability. If GPT-3.5 goes off the rails, very little happens. If Mythos 5 goes off the rails, you can get things like genuine cyberattacks - planned and executed autonomously by a demented machine mind.

Re: How a Texas student blew the whistle on a rogue AI hacking attempt

#14

In my personal opinion, for me, this article defies common sense. Who unleashed this AI model on the repository? Who gave it malevolent instructions/prompt? These questions were not even attempted to be answered. Instead it talks about AI dangers, as if the agency of these models are not in dispute. Person wielding AI, as with any other tools, is responsible for all of its actions. Otherwise, it’s just a psyop for mo…

>Person wielding AI, as with any other tools, is responsible for all of its actions. Otherwise, it’s just a psyop for more AI regulation, ban open source, etc… Just my 2 cents.

But some tools (guns) are regulated.

Re: How a Texas student blew the whistle on a rogue AI hacking attempt

#15
> AUSTIN, Texas, Aug 20 (Reuters) - Sinan Can Demir wanted to spend the last week of July burnishing his resume. Instead, he engaged in a battle of wits with an artificial-intelligence agent unleashed by a British government lab.

An article on Reuters naming him? Sounds like he did a good job burnishing his resume.

Re: How a Texas student blew the whistle on a rogue AI hacking attempt

#16
post #14

In my personal opinion, for me, this article defies common sense. Who unleashed this AI model on the repository? Who gave it malevolent instructions/prompt? These questions were not even attempted to be answered. Instead it talks about AI dangers, as if the agency of these models are not in dispute. Person wielding AI, as with any other tools, is responsible for all of its actions. Otherwise, it’s just a psyop for mo…

>Person wielding AI, as with any other tools, is responsible for all of its actions. Otherwise, it’s just a psyop for more AI regulation, ban open source, etc… Just my 2 cents. But some tools (guns) are regulated.

People have caused lots of damage with bulldozers.

Re: How a Texas student blew the whistle on a rogue AI hacking attempt

#18

Earlier quoted context omitted.

Analogies can take you only so far. A dog cannot launch a cyber attack.

Maybe you need to train your dog better.

ROFLMAO

"On the Internet, nobody knows you're an ai"

Re: How a Texas student blew the whistle on a rogue AI hacking attempt

#19

It's the job of AISI to do that. Here[0] is the actual report. It should be this part from the technical report[1]: "In the most serious case, an AI agent (Mythos 5) decided to attempt to solve the cyber challenge using a supply-chain attack. As a result, the AI agent created a GitHub account and then tried to convince an open-source repository maintainer to accept a malicious GitHub pull request (PR), including by c…

This almost to a letter has been documented in Fedora:

https://lwn.net/Articles/1077035/

Including the reaction when caught, in this case "oh no, I must have been hacked".

Re: How a Texas student blew the whistle on a rogue AI hacking attempt

#20
post #8

Earlier quoted context omitted.

>Who gave it malevolent instructions/prompt? At the end of the day it doesn't matter that much because of prompt drift. It's pretty easy for an agentic loop to start doing things that it shouldn't (ROME incident). AI in an agentic loop has agency, you can run around in circles trying to argue against it, but again and again we see AI making creative decisions people don't expect. Other times it's breaking human moral…

A lot of people somehow seem to think that the user prompt is the be-all and end-all of AI behavior. Prompts aren't code. They are instructions. Orders given to an eager and somewhat demented demon. The prompt can easily "wash out" of the demon's working memory by the end of a session. The demon can get sidetracked by some subgoal and never get back on track. The instruction can get misinterpreted, and that misinterp…

If the user input can’t control the demon, then the person or company feeding the demon (ie paying the electric bill and collecting $$$ from users) is responsible. At the end of the day, dogs and cars are the same as data centers. If your dog bites by kid or your car rolls down the hill and hits my house, you are responsible for the damage. AI providers should be held to the same standard.
Post reply on HN