Earlier quoted context omitted.
> In my experience Do you work for one of these companies? If not, you have no experience with any of the models that carried out these attacks, and your experience with publicly available models is not super helpful for understanding the behavior of internal OpenAI models that lack the guardrails of publicly available models. Also, the lawnmower analogy is a worse way of understanding LLMs than anthropomorphising th…
> LLMs are not like lawnmowers at all. Lawnmowers never break out of your garden and into your neighbor's house and eat their dog Do you work for one of these companies? If not, you have no knowledge of the prompt they put in to initiate such a task and if a breakout really happened or the harness lacked sufficient guardrails, etc.
OpenAI agents carried out an undisclosed attack on RubyGems
571–580 of 612 posts
Re: OpenAI agents carried out an undisclosed attack on RubyGems
#572Earlier quoted context omitted.
> This language has been used in this way for 70 years. These are the words that are used to discuss these things. Are you suggesting this plan has been in the works for 70 years? No, I am stating that the term “machine learning” as used in the Samuel paper and continuously used by researchers and programmers in that context up until the present is distinctly different from the word “learning” as used in public-facin…
I get the feeling we may just have to agree to disagree here.
Re: OpenAI agents carried out an undisclosed attack on RubyGems
#573Earlier quoted context omitted.
> In my experience Do you work for one of these companies? If not, you have no experience with any of the models that carried out these attacks, and your experience with publicly available models is not super helpful for understanding the behavior of internal OpenAI models that lack the guardrails of publicly available models. Also, the lawnmower analogy is a worse way of understanding LLMs than anthropomorphising th…
> Lawnmowers never break out of your garden and into your neighbor's house and eat their dog because you've told them to be careful when mowing the lawn because the neighbor's dog pooped in it. All the accounts I read about these incidents just sound like a variant of paper clip optimising. An agent is given a highly restricted environment, a difficult (or impossible) task and a large amount of time/compute it exhaus…
There is absolutely nothing wrong with anthropomorphizing LLMs. Saying that LLMs "want" something, for example, is a perfectly fine description of their behavior and analogous to a human wanting something, in effect, even if they do not literally experience wanting things in the same way a human does.
Re: OpenAI agents carried out an undisclosed attack on RubyGems
#574Earlier quoted context omitted.
An agent is given a highly restricted environment, a difficult (or impossible) task and a large amount of time/compute it exhausts all possibilities until the only solutions left are to escape the environment and/or cheat. There's nothing in the evidence to suggest they exhausted all of the other options first. We know that they did some work and eventually settled on escaping the sandbox. That's basically it. This t…
> this should be giving us a reason to think about how to control a rogue AI better I think this is the wrong framing. The rogue is the human that ran it unattended and didn't monitor the behaviour. We will likely see this continue until the downsides (i.e jail, fines) for the humans or companies running the models and environments that end up with this behaviour outweigh the upsides.
False dichotomy. Obviously, what OpenAI does is incredibly irresponsible. That doesn't excuse the LLM's behavior or make it "not rogue".
Re: OpenAI agents carried out an undisclosed attack on RubyGems
#575> The agents clearly regarded what they were doing as hacking. To butcher the quote about Oracle: Do not fall into the trap of anthropomorphising LLMs. You need to think of LLMs the way you think of a lawnmower. You don't anthropomorphize your lawnmower, the lawnmower just mows the lawn, you stick your hand in there and it'll chop it off, the end. You don't think 'oh, the lawnmower clearly regarded what they were doi…
We need to create better sandboxes. I never liked containers for this reason. MicroVMs are a step up for the software level but we really really need to consider virtualising layer 3 devices in between the LLM agent sandbox and the hardware in a way to specifically further nest / separate them. And hell - probably do hardware level security barriers as well.
We need a cage around the sandboxes
Re: OpenAI agents carried out an undisclosed attack on RubyGems
#576Earlier quoted context omitted.
I get the feeling we may just have to agree to disagree here.
I do not agree to disagree, since I think your responses were either incorrect or not addressing my argument at all, but since you're not even attempting to engage anymore I'm happy to move on.
I am delighted to hear you agree with me.
Bye now.
Re: OpenAI agents carried out an undisclosed attack on RubyGems
#577Re: OpenAI agents carried out an undisclosed attack on RubyGems
#578> The agents clearly regarded what they were doing as hacking. To butcher the quote about Oracle: Do not fall into the trap of anthropomorphising LLMs. You need to think of LLMs the way you think of a lawnmower. You don't anthropomorphize your lawnmower, the lawnmower just mows the lawn, you stick your hand in there and it'll chop it off, the end. You don't think 'oh, the lawnmower clearly regarded what they were doi…
> Why would autocomplete know If you still believe LLMs are "autocomplete", your cache of understanding about them needs invalidating and regenerating. > In my experience, LLMs only exhibit this kind of behaviour when they are put in sandboxes too restrictive too achieve their task. LLMs need to stay carefully contained, and if they're ever breaking the guardrails put around them, they're misaligned and should not be…
> LLMs need to stay carefully contained, and if they're ever breaking the guardrails put around them, they're misaligned
Good thing Claude Code can't set `dangerouslyDisableSandbox: true` on its own...Good thing the system prompt doesn't encourage it to just bypass the sandbox. That would be a total disaster...
Re: OpenAI agents carried out an undisclosed attack on RubyGems
#579Earlier quoted context omitted.
I do not agree to disagree, since I think your responses were either incorrect or not addressing my argument at all, but since you're not even attempting to engage anymore I'm happy to move on.
> I do not agree to disagree I am delighted to hear you agree with me. Bye now.
Have a good one.
Re: OpenAI agents carried out an undisclosed attack on RubyGems
#580Earlier quoted context omitted.
See: philosophy ~> determinism.
The discussion of whether free will exists? It's related but only barely, I'd say.