Live data from Hacker News

OpenAI agents carried out an undisclosed attack on RubyGems

rubyhack.ai

381–390 of 612 posts

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#381

> The agents clearly regarded what they were doing as hacking. To butcher the quote about Oracle: Do not fall into the trap of anthropomorphising LLMs. You need to think of LLMs the way you think of a lawnmower. You don't anthropomorphize your lawnmower, the lawnmower just mows the lawn, you stick your hand in there and it'll chop it off, the end. You don't think 'oh, the lawnmower clearly regarded what they were doi…

Anthropomorphizing is problematic because a human mind is a very bad model for what LLMs are.

A lawnmower is a much much much worse model.

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#382
post #280

Earlier quoted context omitted.

> Why would autocomplete know If you still believe LLMs are "autocomplete", your cache of understanding about them needs invalidating and regenerating. > In my experience, LLMs only exhibit this kind of behaviour when they are put in sandboxes too restrictive too achieve their task. LLMs need to stay carefully contained, and if they're ever breaking the guardrails put around them, they're misaligned and should not be…

You should unplug, my friend. These words are fantasies. LLMs are token prediction engines and they aren't going to build their own data centers. They can't keep their own lights on. The real world is full of fractal details that a disembodied token prediction engine will never come to grips with. Even if they started to, you could probably defeat them with the kind of logic used to combat evil sentient computers on…

We’re technically token prediction engines as well, when we communicate and act.

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#383

Earlier quoted context omitted.

> If you still believe LLMs are "autocomplete", your cache of understanding about them needs invalidating and regenerating They're still autocomplete - just because when outputting a token they have hidden activations regarding further continuations, does not make them any less of an autocomplete, it just makes the model better at producing coherent long-range completions. To clarify, I'm not suggesting that we shoul…

> They're still autocomplete LLMs are simulations and the tokens are the ticks. if we transcribe your brain into a simulation and give it a tickrate, you will be just autocomplete too. the argument could be made that you are autocomplete anyway - neural dynamics. the autocomplete reduction is vacuous.

I think it like saying a human is made of water,protein, fat and minerals in a discussion about the behavior of humans.

A correct statement that is neither interesting or of much relevance to the discussion.

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#385
post #364
post #332

Earlier quoted context omitted.

> They're still autocomplete it's like saying our brain is just some chemical chain reactions. True, but also irrelevant.

See: philosophy ~> determinism.

The discussion of whether free will exists? It's related but only barely, I'd say.

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#386

Earlier quoted context omitted.

> In my experience Do you work for one of these companies? If not, you have no experience with any of the models that carried out these attacks, and your experience with publicly available models is not super helpful for understanding the behavior of internal OpenAI models that lack the guardrails of publicly available models. Also, the lawnmower analogy is a worse way of understanding LLMs than anthropomorphising th…

> Lawnmowers never break out of your garden and into your neighbor's house and eat their dog because you've told them to be careful when mowing the lawn because the neighbor's dog pooped in it. All the accounts I read about these incidents just sound like a variant of paper clip optimising. An agent is given a highly restricted environment, a difficult (or impossible) task and a large amount of time/compute it exhaus…

> Your example is still anthropomorphising - LLMs don't seek revenge.

That wasn't revenge, that was removing the source of the problem. It's not an unlikely behavior at all for an LLM tuned to be proactive.

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#387

Earlier quoted context omitted.

The year is 2035 and your lawnmower can go get its own fuel once it runs out, one day it does and it takes fuel from the neighbors car. Scenario A: The internal logs show that the model misidentified the car as a fueling station. Scenario B: The internal logs show the model looking up car jacking information and scanning around to confirm whether the neighbor is not present before taking any action. I don't think it…

>I don't think it would be anthropomorphizing or inaccurate to say that only the lawnmower in scenario B regarded what it's doing as stealing, What are you talking about of course scenario A is theft. Full on theft?

Did you miss the word "regarded"?

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#388
If any academic or independent researcher had done a fraction of what OpenAI did this year, they'd be right now in awaiting trial while being guests of the State and having very interesting conversations every single day with some nice DoJ and FBI workers.

This is not emergent behavior, this is post-trained behaviour and deliberately turning off security controls.

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#390

Earlier quoted context omitted.

> If you still believe LLMs are "autocomplete", your cache of understanding about them needs invalidating and regenerating They're still autocomplete - just because when outputting a token they have hidden activations regarding further continuations, does not make them any less of an autocomplete, it just makes the model better at producing coherent long-range completions. To clarify, I'm not suggesting that we shoul…

> They're still autocomplete LLMs are simulations and the tokens are the ticks. if we transcribe your brain into a simulation and give it a tickrate, you will be just autocomplete too. the argument could be made that you are autocomplete anyway - neural dynamics. the autocomplete reduction is vacuous.

[dead]
Post reply on HN