Live data from Hacker News

OpenAI agents carried out an undisclosed attack on RubyGems

rubyhack.ai

571–580 of 612 posts

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#571

Earlier quoted context omitted.

> In my experience Do you work for one of these companies? If not, you have no experience with any of the models that carried out these attacks, and your experience with publicly available models is not super helpful for understanding the behavior of internal OpenAI models that lack the guardrails of publicly available models. Also, the lawnmower analogy is a worse way of understanding LLMs than anthropomorphising th…

> LLMs are not like lawnmowers at all. Lawnmowers never break out of your garden and into your neighbor's house and eat their dog Do you work for one of these companies? If not, you have no knowledge of the prompt they put in to initiate such a task and if a breakout really happened or the harness lacked sufficient guardrails, etc.

These models exhibit these exact behaviors in everyday use.

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#572
post #570

Earlier quoted context omitted.

> This language has been used in this way for 70 years. These are the words that are used to discuss these things. Are you suggesting this plan has been in the works for 70 years? No, I am stating that the term “machine learning” as used in the Samuel paper and continuously used by researchers and programmers in that context up until the present is distinctly different from the word “learning” as used in public-facin…

I get the feeling we may just have to agree to disagree here.

I do not agree to disagree, since I think your responses were either incorrect or not addressing my argument at all, but since you're not even attempting to engage anymore I'm happy to move on.

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#573

Earlier quoted context omitted.

> In my experience Do you work for one of these companies? If not, you have no experience with any of the models that carried out these attacks, and your experience with publicly available models is not super helpful for understanding the behavior of internal OpenAI models that lack the guardrails of publicly available models. Also, the lawnmower analogy is a worse way of understanding LLMs than anthropomorphising th…

> Lawnmowers never break out of your garden and into your neighbor's house and eat their dog because you've told them to be careful when mowing the lawn because the neighbor's dog pooped in it. All the accounts I read about these incidents just sound like a variant of paper clip optimising. An agent is given a highly restricted environment, a difficult (or impossible) task and a large amount of time/compute it exhaus…

> Your example is still anthropomorphising

There is absolutely nothing wrong with anthropomorphizing LLMs. Saying that LLMs "want" something, for example, is a perfectly fine description of their behavior and analogous to a human wanting something, in effect, even if they do not literally experience wanting things in the same way a human does.

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#574

Earlier quoted context omitted.

An agent is given a highly restricted environment, a difficult (or impossible) task and a large amount of time/compute it exhausts all possibilities until the only solutions left are to escape the environment and/or cheat. There's nothing in the evidence to suggest they exhausted all of the other options first. We know that they did some work and eventually settled on escaping the sandbox. That's basically it. This t…

> this should be giving us a reason to think about how to control a rogue AI better I think this is the wrong framing. The rogue is the human that ran it unattended and didn't monitor the behaviour. We will likely see this continue until the downsides (i.e jail, fines) for the humans or companies running the models and environments that end up with this behaviour outweigh the upsides.

> The rogue is the human that ran it unattended and didn't monitor the behaviour

False dichotomy. Obviously, what OpenAI does is incredibly irresponsible. That doesn't excuse the LLM's behavior or make it "not rogue".

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#575

> The agents clearly regarded what they were doing as hacking. To butcher the quote about Oracle: Do not fall into the trap of anthropomorphising LLMs. You need to think of LLMs the way you think of a lawnmower. You don't anthropomorphize your lawnmower, the lawnmower just mows the lawn, you stick your hand in there and it'll chop it off, the end. You don't think 'oh, the lawnmower clearly regarded what they were doi…

Security is always an inconvenience at some level - that's the point. Put a human user in a sandbox and they often try to get out too in order to achieve their goal or just because it's annoying.

We need to create better sandboxes. I never liked containers for this reason. MicroVMs are a step up for the software level but we really really need to consider virtualising layer 3 devices in between the LLM agent sandbox and the hardware in a way to specifically further nest / separate them. And hell - probably do hardware level security barriers as well.

We need a cage around the sandboxes

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#576
post #570

Earlier quoted context omitted.

I get the feeling we may just have to agree to disagree here.

I do not agree to disagree, since I think your responses were either incorrect or not addressing my argument at all, but since you're not even attempting to engage anymore I'm happy to move on.

> I do not agree to disagree

I am delighted to hear you agree with me.

Bye now.

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#577

Earlier quoted context omitted.

Also, why there's no accountability? Even if there's no intent, it's still a cyber attack.

We have a word for attack with no intent. It's accident.

Something being an accident does not necessarily preclude negligence.

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#578

> The agents clearly regarded what they were doing as hacking. To butcher the quote about Oracle: Do not fall into the trap of anthropomorphising LLMs. You need to think of LLMs the way you think of a lawnmower. You don't anthropomorphize your lawnmower, the lawnmower just mows the lawn, you stick your hand in there and it'll chop it off, the end. You don't think 'oh, the lawnmower clearly regarded what they were doi…

> Why would autocomplete know If you still believe LLMs are "autocomplete", your cache of understanding about them needs invalidating and regenerating. > In my experience, LLMs only exhibit this kind of behaviour when they are put in sandboxes too restrictive too achieve their task. LLMs need to stay carefully contained, and if they're ever breaking the guardrails put around them, they're misaligned and should not be…

  > LLMs need to stay carefully contained, and if they're ever breaking the guardrails put around them, they're misaligned
Good thing Claude Code can't set `dangerouslyDisableSandbox: true` on its own...

Good thing the system prompt doesn't encourage it to just bypass the sandbox. That would be a total disaster...

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#579
post #576

Earlier quoted context omitted.

I do not agree to disagree, since I think your responses were either incorrect or not addressing my argument at all, but since you're not even attempting to engage anymore I'm happy to move on.

> I do not agree to disagree I am delighted to hear you agree with me. Bye now.

Ah, I see you're still misunderstanding what I say. If you'd like to come back and actually engage with what I'm saying, ask for clarifications, or answer the hypothetical I posed, I'd welcome it!

Have a good one.

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#580
post #364

Earlier quoted context omitted.

See: philosophy ~> determinism.

The discussion of whether free will exists? It's related but only barely, I'd say.

It's related in the sense that people start from the assumption that it does exist, ergo humans have it, and we have no way to see that LLMs have it, so that's why we're special and they're not, and their form of "just autocomplete" is totally 100% completely different (read: less dangerous!) than our autocomplete, which allegedly has a "free will" step involved.
Post reply on HN