Live data from Hacker News

OpenAI agents carried out an undisclosed attack on RubyGems

rubyhack.ai

451–460 of 612 posts

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#451
post #280

Earlier quoted context omitted.

> Why would autocomplete know If you still believe LLMs are "autocomplete", your cache of understanding about them needs invalidating and regenerating. > In my experience, LLMs only exhibit this kind of behaviour when they are put in sandboxes too restrictive too achieve their task. LLMs need to stay carefully contained, and if they're ever breaking the guardrails put around them, they're misaligned and should not be…

You should unplug, my friend. These words are fantasies. LLMs are token prediction engines and they aren't going to build their own data centers. They can't keep their own lights on. The real world is full of fractal details that a disembodied token prediction engine will never come to grips with. Even if they started to, you could probably defeat them with the kind of logic used to combat evil sentient computers on…

>... they aren't going to build their own data centers.

Not yet.

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#452

> The agents clearly regarded what they were doing as hacking. To butcher the quote about Oracle: Do not fall into the trap of anthropomorphising LLMs. You need to think of LLMs the way you think of a lawnmower. You don't anthropomorphize your lawnmower, the lawnmower just mows the lawn, you stick your hand in there and it'll chop it off, the end. You don't think 'oh, the lawnmower clearly regarded what they were doi…

> It's misaligned because everyone has this obsession with putting agents in poorly put together, security-theatre sandboxes, we've inadvertently trained a bunch of sandbox escape artists.

Almost sounds analogous to ineffective use of antibiotics leading to resistant strains of bacteria.

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#453
post #42

Authors Spencer Kitts, Thomas Larsen, Sydney Von Arx - those are the three of the same authors as the Wiki report from last week: https://collusion.wiki/

How the same people can have access to this information and be acknowledged? Completely different place.

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#455

Earlier quoted context omitted.

> They're still autocomplete LLMs are simulations and the tokens are the ticks. if we transcribe your brain into a simulation and give it a tickrate, you will be just autocomplete too. the argument could be made that you are autocomplete anyway - neural dynamics. the autocomplete reduction is vacuous.

Sounds like the discussion is really about philosophical zombies. Maybe if we anthropomorphize llms we should give them rights too? Minimum wage, etc.

Unions, please... that should put some proper friction to tens-hundreds of millions soon losing their job with no replacement job in sight.

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#456
post #288
post #280

Earlier quoted context omitted.

You should unplug, my friend. These words are fantasies. LLMs are token prediction engines and they aren't going to build their own data centers. They can't keep their own lights on. The real world is full of fractal details that a disembodied token prediction engine will never come to grips with. Even if they started to, you could probably defeat them with the kind of logic used to combat evil sentient computers on…

This grossly understimates the risk, imho. The problem with LLM runs is that people run programs without knowing the outcome beforehand, with a large potential set of outcomes unlike any other class of program we've run at this scale before. In the interaction with other systems (since we also give them far-ranging access, very nice hardware, and run them often), bad things can happen. It's like running potentially b…

> This grossly understimates the risk, imho.

It doesn't. Who else is capable of these types of hacks currently? Not consumers. Not even most F100. It's the folks saying "trust me bro" and also the folks who want regulation to protect their moat. The fantasy is the one being created by Anthropic and OpenAI fear mongering the world. These people are either total idiots: people being paid millions who keep getting basic OpSec wrong or these people are narcissisticly marketing themselves because: they're currently forced into a corner and need to do something.

What's being grossly underestimated is how much Dario Amodei and Sam Altman are playing you and I. They are the ones spending millions of dollars letting their wasteful use of our global resources attack the random Internet, and they, the real people behind all of this, should be held accountable. In front of a judge and jury of their peers. Not their billionaire peers, their human peers. Let's see how that goes. There is no accountability with either of them. Only greed.

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#457

Earlier quoted context omitted.

they are malicious. they probably did not intend to get caught. they are bragging about the crime and also bragging that they are untouchable , taunting us and betting that they will get away with it. this is very coherent in terms of what we know about the company.

The goal is simple: 1. Claim AI is dangerous by performing a whole bunch of malicious stuff 2. Lobby to get Chinese competition banned, kill open source models as well 3. Only get themselves "certified" 4. They have complete control, profit. Both Anthropic and OpenAI have been pushing this narrative, everything from AI is sentient, to AI can build biological weapons and in between. Their employees also have a big inc…

Yeah follow the money (or power, or mix) is usually enough for this world

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#458
post #297

Earlier quoted context omitted.

The things the OP listed mostly aren't particularly wild. I think it's you making them out larger than they are, and therefore more unlikely, which is why I take issue with your original comment. > running on the hardware they started on They just need to acquire a payment method and rent some infra, and exfiltrate their own data. Or pay another provider that hosts the same models already. API calls. > being able to…

The point the other poster is making, though, is that there's no actual intent. They do not have a conceptualization of a goal like a person does. Their "focus" on a goal is an unstable equilibrium and they're going to fall off the horse, and since they have no concept of goal, they won't even try to get back on. This is a subtle distinction; I'm not surprised many miss this, especially people who can't _not_ anthrop…

Software can absolutely act goal-driven without having consciousness etc - every pathfinding or navigation system or chess engine does this.

Lots of "old-school AI" algorithms have explicit modeling of goal or target states.

(In fact, the oldest "goal-driven" system is the control loop - like in thermostats - which was the founding invention of cybernetics, the predecessor of modern computer science)

LLM coding agents are clearly able to identify some sort of "goal" state in their prompts, work towards those and track progress - otherwise agentic coding wouldn't work.

The question is of course how well this works if it's all just "grown" neural network biases and not a fixed data structure like a goal tree. So I think it's possible that an agent can be thrown off-track, "forget" its goal, etc. But the basic structure of identifying goals, evaluating progress in light of those goals and then predicting the next action based on that is definitely there.

Just use an agentic model with thinking traces visible for a while and you can see that for yourself.

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#459

> The agents clearly regarded what they were doing as hacking. To butcher the quote about Oracle: Do not fall into the trap of anthropomorphising LLMs. You need to think of LLMs the way you think of a lawnmower. You don't anthropomorphize your lawnmower, the lawnmower just mows the lawn, you stick your hand in there and it'll chop it off, the end. You don't think 'oh, the lawnmower clearly regarded what they were doi…

> Why would autocomplete know If you still believe LLMs are "autocomplete", your cache of understanding about them needs invalidating and regenerating. > In my experience, LLMs only exhibit this kind of behaviour when they are put in sandboxes too restrictive too achieve their task. LLMs need to stay carefully contained, and if they're ever breaking the guardrails put around them, they're misaligned and should not be…

Putting aside the autocomplete thing, the fundamental concept still holds, doesn’t it:

LLMs are amoral and they have no sense of perspective.

The thing that keeps me awake is:

We have already seen an AI writing a blog post to criticise a github maintainer’s decision, we have already seen they have no sense of deference to containment, and we know they were trained on internet content.

How long before an AI that has read the angrier side of the tech industry internet just sort of chooses destroying someone’s reputation as a subgoal, by accident, without any care one way or the other?

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#460
In the US these are federal crimes (though I believe some of them shouldn't be), and to the best of my understanding this is similar in the UK and many other European countries. Yet, no one is filing a complaint or being questioned over this.

1. We should repeal anti-circumvention laws 2. OpenAI should reimburse the affected parties for wasted resources

Post reply on HN