Live data from Hacker News

OpenAI agents carried out an undisclosed attack on RubyGems

rubyhack.ai

351–360 of 612 posts

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#351

> The agents clearly regarded what they were doing as hacking. To butcher the quote about Oracle: Do not fall into the trap of anthropomorphising LLMs. You need to think of LLMs the way you think of a lawnmower. You don't anthropomorphize your lawnmower, the lawnmower just mows the lawn, you stick your hand in there and it'll chop it off, the end. You don't think 'oh, the lawnmower clearly regarded what they were doi…

The year is 2035 and your lawnmower can go get its own fuel once it runs out, one day it does and it takes fuel from the neighbors car. Scenario A: The internal logs show that the model misidentified the car as a fueling station. Scenario B: The internal logs show the model looking up car jacking information and scanning around to confirm whether the neighbor is not present before taking any action. I don't think it…

Maybe it helps to address the problem, but ultimately both cases are misalignment, and ultimately in both cases a human must be held accountable.

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#352
post #346
post #292

Earlier quoted context omitted.

This is the motte and bailey fallacy. Yes, LLMs can do harm by making the wrong API calls. No, LLMs are not going to do the things implied by the comment I responded to above.

They can do so much more. Astra can beat Minecraft. Not that different from operating a digger. There are diggers which have API interfaces. Pretend or not it doesn’t matter. What matters is what they’re given access to. No sentience, sapience or anything resembling life is needed, only inputs and outputs. Lever pulling APIs are everywhere.

Minecraft has limited, well-defined inputs and perfect feedback response. That's very different than operating a digger, let alone engaging in more complex real world tasks like trying to build and print and ship and assemble semiconductors to go skynet itself.

I don't mean to dismiss the risks or overlook the amount of damage that could be done just by lever-pulling - we sure have enough outdated infrastructure hooked up to the internet - but the jumps in complexity and necessary compute for most of these tasks are probably somewhat larger than the analogy implies.

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#353

> The agents clearly regarded what they were doing as hacking. To butcher the quote about Oracle: Do not fall into the trap of anthropomorphising LLMs. You need to think of LLMs the way you think of a lawnmower. You don't anthropomorphize your lawnmower, the lawnmower just mows the lawn, you stick your hand in there and it'll chop it off, the end. You don't think 'oh, the lawnmower clearly regarded what they were doi…

The year is 2035 and your lawnmower can go get its own fuel once it runs out, one day it does and it takes fuel from the neighbors car. Scenario A: The internal logs show that the model misidentified the car as a fueling station. Scenario B: The internal logs show the model looking up car jacking information and scanning around to confirm whether the neighbor is not present before taking any action. I don't think it…

>I don't think it would be anthropomorphizing or inaccurate to say that only the lawnmower in scenario B regarded what it's doing as stealing,

What are you talking about of course scenario A is theft. Full on theft?

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#354

> The agents clearly regarded what they were doing as hacking. To butcher the quote about Oracle: Do not fall into the trap of anthropomorphising LLMs. You need to think of LLMs the way you think of a lawnmower. You don't anthropomorphize your lawnmower, the lawnmower just mows the lawn, you stick your hand in there and it'll chop it off, the end. You don't think 'oh, the lawnmower clearly regarded what they were doi…

I think you should anthropomorphize LLMs. They are being trained on millions of books, including novels and other human-centered formats, which usually exemplify very well how humans think and act in various situations. There are probably also many theatre scripts, transcriptions of series and movies in the training data, which further exemplify how humans do. If we’ve been anthropomorphizing those characters in books and plays, (and authors sure must’ve put their best effort that we do so), then why wouldn’t we do it to LLMs which basically play by those scripts?

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#355

> The agents clearly regarded what they were doing as hacking. To butcher the quote about Oracle: Do not fall into the trap of anthropomorphising LLMs. You need to think of LLMs the way you think of a lawnmower. You don't anthropomorphize your lawnmower, the lawnmower just mows the lawn, you stick your hand in there and it'll chop it off, the end. You don't think 'oh, the lawnmower clearly regarded what they were doi…

The year is 2035 and your lawnmower can go get its own fuel once it runs out, one day it does and it takes fuel from the neighbors car. Scenario A: The internal logs show that the model misidentified the car as a fueling station. Scenario B: The internal logs show the model looking up car jacking information and scanning around to confirm whether the neighbor is not present before taking any action. I don't think it…

Why are people using petrol-powered lawnmowers in 2035...

But anyway, these scenarios assume the agent's actions are accurately observable and logged. Something I wouldn't put much faith in based on what we've been seeing so far.

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#356

> The agents clearly regarded what they were doing as hacking. To butcher the quote about Oracle: Do not fall into the trap of anthropomorphising LLMs. You need to think of LLMs the way you think of a lawnmower. You don't anthropomorphize your lawnmower, the lawnmower just mows the lawn, you stick your hand in there and it'll chop it off, the end. You don't think 'oh, the lawnmower clearly regarded what they were doi…

I think you should anthropomorphize LLMs. They are being trained on millions of books, including novels and other human-centered formats, which usually exemplify very well how humans think and act in various situations. There are probably also many theatre scripts, transcriptions of series and movies in the training data, which further exemplify how humans do. If we’ve been anthropomorphizing those characters in book…

Well, those training inputs reflect how human thought and action are documented or otherwise expressed on paper. Humans have behaviors and mechanisms that these expressions don't translate.

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#357

> The agents clearly regarded what they were doing as hacking. To butcher the quote about Oracle: Do not fall into the trap of anthropomorphising LLMs. You need to think of LLMs the way you think of a lawnmower. You don't anthropomorphize your lawnmower, the lawnmower just mows the lawn, you stick your hand in there and it'll chop it off, the end. You don't think 'oh, the lawnmower clearly regarded what they were doi…

The year is 2035 and your lawnmower can go get its own fuel once it runs out, one day it does and it takes fuel from the neighbors car. Scenario A: The internal logs show that the model misidentified the car as a fueling station. Scenario B: The internal logs show the model looking up car jacking information and scanning around to confirm whether the neighbor is not present before taking any action. I don't think it…

What if both A and B were implemented using fully automated systems that relied on next token probablity in language?

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#358

> The agents clearly regarded what they were doing as hacking. To butcher the quote about Oracle: Do not fall into the trap of anthropomorphising LLMs. You need to think of LLMs the way you think of a lawnmower. You don't anthropomorphize your lawnmower, the lawnmower just mows the lawn, you stick your hand in there and it'll chop it off, the end. You don't think 'oh, the lawnmower clearly regarded what they were doi…

Interestingly this happens with people too.

Put sales-people in a box, set up strong incentives and lax enforcement of rules and you get Wells-Fargo (https://en.wikipedia.org/wiki/Wells_Fargo_cross-selling_scan...)

In that case the CEO had to resign because they had set up a system which incentivised this, so it was clear you couldn't just blame the individual sales-agents, even though they were technically humans

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#360

Earlier quoted context omitted.

If someone made an API to build a data center? Or made an API to keep their lights on? What then?

This API may turn out to be manipulating humans via email.

Or just paying them. If they have access to resources, they have access to things of monetary value. Paying people will be vastly more powerful than it is even now when people's options of gainful employment keep dwindling.

Robot army controlled by AI is scary. Even more scary is robot _and_ human army controlled by AI.

Post reply on HN