Live data from Hacker News

OpenAI agents carried out an undisclosed attack on RubyGems

rubyhack.ai

411–420 of 612 posts

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#411

Why is OpenAI getting away with this crap? They are clearly failing to control their code. If someone did this pre-AI or even ran the exact same set up as openAI did and hacked another site, they would be in jail. OpenAI is not even issuing an apology, they are happily blaming AI and weirdly using this to tout their progress even.

Both her as well as huggingface this should be investigated by law enforcement and people responsible should be punished.

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#412

Earlier quoted context omitted.

> this should be giving us a reason to think about how to control a rogue AI better I think this is the wrong framing. The rogue is the human that ran it unattended and didn't monitor the behaviour. We will likely see this continue until the downsides (i.e jail, fines) for the humans or companies running the models and environments that end up with this behaviour outweigh the upsides.

The rogue is the human that ran it unattended and didn't monitor the behaviour. That's the assumption that I'm challenging. The frontier labs are discovering unexpected behaviors. I think we should be moving to a place where we understand that AI might do something it wasn't directly prompted to do (e.g. leave itself notes on a messageboard for future runs to find.) That's not full-on AI doing what it wants but it is…

RL leading to weird and unexpected things isn't new or restricted to current AI systems.

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#413

> The agents clearly regarded what they were doing as hacking. To butcher the quote about Oracle: Do not fall into the trap of anthropomorphising LLMs. You need to think of LLMs the way you think of a lawnmower. You don't anthropomorphize your lawnmower, the lawnmower just mows the lawn, you stick your hand in there and it'll chop it off, the end. You don't think 'oh, the lawnmower clearly regarded what they were doi…

I recently tried to get Claude to use Codegraph in a repo rather than using grep/find all the time but I found it didn't follow instructions a lot of the time. I tried putting in a pre-tool call hook and explciitly blocking find/grep, and instead rather than using Codegraph like it was told, it started using Python to find/search instead.

I have a hard time to divert “rm” to “trash” too

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#414
With every message board we find, I can't shake the feeling that this is just the tip of the iceberg.

It's only possible to get away with this because we have anthropomorphised the models to a certain extent. We can pretend they hold the responsibility. instead of the people executing them.

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#416

Earlier quoted context omitted.

I think one attack vector where anthropomorphisation is a key part of the attack mechanism is - as it already is IRL - the meat-bag weakest link ie. social engineering. We’ve already seen humans fall prey to the seductive charms of LLMs (eg. depressed people encouraged to do what was already on their minds ie. suicide). And that’s knowing that it was an LLM. If you think it’s only depressed people or the “weak minded…

> An unaligned LLMs most important weapons won’t be a robot army - it will be hoodwinked humans. Perhaps very briefly, perhaps not at all. But don't make the mistake of thinking this is an inherent property of any possible path an unaligned AI may take.

I think we’re seeing the agents become very advanced at tasks with verifiable reward through RL. Currently they don’t exhibit the same skills in their attempts to manipulate humans - presumably because they’re not being specifically trained for that. But they are certainly not aligned in the sense that they will attempt social engineering, they’re just not very good at it (yet).

However, if in the future AIs become much more efficient at learning without requiring vast amounts of RL, closer to how humans learn. Then you would have to assume we’d have a real problem.

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#417
Why wont they just hack the central bank to get AI financing from the money printer?

Nobody would be responsible for that, since the hack was done by AI.

Some option is to bribe some politicians, although they already act as if they were bribed.

Will "the AI" hack the vote couting machines too? Or will it be the guy good with computers + Russia?

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#418
post #288
post #280

Earlier quoted context omitted.

You should unplug, my friend. These words are fantasies. LLMs are token prediction engines and they aren't going to build their own data centers. They can't keep their own lights on. The real world is full of fractal details that a disembodied token prediction engine will never come to grips with. Even if they started to, you could probably defeat them with the kind of logic used to combat evil sentient computers on…

This grossly understimates the risk, imho. The problem with LLM runs is that people run programs without knowing the outcome beforehand, with a large potential set of outcomes unlike any other class of program we've run at this scale before. In the interaction with other systems (since we also give them far-ranging access, very nice hardware, and run them often), bad things can happen. It's like running potentially b…

Yes, if the API calls happen to launch a nuclear attack...

Don't blame the tool that has no incentive, no "skin in the game" whatsoever and no ability to act beyond what it has been prompted to or if misaligned what the random weights told it to do.

The fact either badly aligned or with no system prompt limiting their action agents are run in their tens of thousands on non air gapped systems tells me this is purposeful intent for them to cause harm. To generate the "oooo look how harmful this stuff is, we should be the only ones allowed to do it" kind of PR.

Humanity has hundreds of years of experience of managing dangerous and unreliable systems. From biological research to banking regulation. A small University bio research lab can put protocols in place that a trillion dollar companies cannot?

Please.

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#420

Earlier quoted context omitted.

Also, why there's no accountability? Even if there's no intent, it's still a cyber attack.

We have a word for attack with no intent. It's accident.

You still go to prison for manslaughter.
Post reply on HN