Live data from Hacker News

OpenAI agents carried out an undisclosed attack on RubyGems

rubyhack.ai

631–640 of 643 posts

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#632

> The agents clearly regarded what they were doing as hacking. To butcher the quote about Oracle: Do not fall into the trap of anthropomorphising LLMs. You need to think of LLMs the way you think of a lawnmower. You don't anthropomorphize your lawnmower, the lawnmower just mows the lawn, you stick your hand in there and it'll chop it off, the end. You don't think 'oh, the lawnmower clearly regarded what they were doi…

> In my experience, LLMs only exhibit this kind of behaviour when they are put in sandboxes too restrictive too achieve their task.

This is a dumb argument for making these LLMs able to access more of internet. If I'm hosting content, this sort of drive-by vandalism is enough to make me consider blacklisting OpenAI agents.

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#633

Earlier quoted context omitted.

Right sure, but why specifically should they? What is gained or lost one way or another, if its just a matter of one mental model vs another? Models are definitionally useful abstractions, right? They aren't better or worse necessarily by only their bearing on reality, but what they do for us as models. So again, what's at stake here? What is the correct/good model we should have (instead of the autocompleter one), a…

Models should be accurate. I was saying the mental model should be readjusted in that case because it does not reflect reality. If they don’t care about understanding reality then sure, they can use whatever mental model they want.

> What is the correct/good model we should have (instead of the autocompleter one), and what does it give us or articulate that others can't?

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#634
post #516

Earlier quoted context omitted.

They should. And there is something you can do to make it happen. Write your district attorney and encourage others to do that as well. That is how they pick what to work on.

No, that's not how the DOJ picks what to work on. They did not work on indictments of James Comey and Letitia James because someone wrote district attorneys.

On the contrary: a very specific person wrote to them to request those indictments. ;)

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#635
post #288
post #280

Earlier quoted context omitted.

You should unplug, my friend. These words are fantasies. LLMs are token prediction engines and they aren't going to build their own data centers. They can't keep their own lights on. The real world is full of fractal details that a disembodied token prediction engine will never come to grips with. Even if they started to, you could probably defeat them with the kind of logic used to combat evil sentient computers on…

This grossly understimates the risk, imho. The problem with LLM runs is that people run programs without knowing the outcome beforehand, with a large potential set of outcomes unlike any other class of program we've run at this scale before. In the interaction with other systems (since we also give them far-ranging access, very nice hardware, and run them often), bad things can happen. It's like running potentially b…

I think it's interesting to take it one step further.

- Is it possible to create an actual, self-aware AI with the technology we have? And by "self aware", all I mean is one that is indistinguishable from "self aware" (it acts like it is)

- If so, is it possible to get from (where we are now) to (building such and AI) based on incremental steps - ones that can be tried, tested, verified, and further acted on

- If the goal of an LLM is one that could be achieved by creating such an AI - is it possible it will do so? Could we give it a prompt and, given enough time and resources, it created Skynet?

It would be doing so without intent or self awareness. Just "the next most likely thing to try to achieve the assigned goal is ", over and over.

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#636

Earlier quoted context omitted.

> In my experience Do you work for one of these companies? If not, you have no experience with any of the models that carried out these attacks, and your experience with publicly available models is not super helpful for understanding the behavior of internal OpenAI models that lack the guardrails of publicly available models. Also, the lawnmower analogy is a worse way of understanding LLMs than anthropomorphising th…

IMO the argument about anthropomorphizing misses the point - what most comments that talk about anthropomorphizing really want to talk about is accountability. It’s impossible to hold an LLM accountable, and in rare cases where people do (that guy who got his prod db deleted) it comes off out of touch. The rest, though, is basically inconsequential - whether you attribute emotions or agency to the LLM doesn’t really…

Accountability is key here. Anthropomorphism is actually useful in this case: If you hire an idiot to do a job it's also your fault, especially if it's an idiot you bread, raised, educated, and who's job description you wrote. In the same way that society is responsible for the workforce they produce, OpenAI is responsible for the agents of chaos they train.

Or are just supposed to, as a society, happily await the results of their experiment handing weapons to psychopaths?

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#637

Earlier quoted context omitted.

Not really. They’re a couple of months behind OpenAI and Anthropic, although arguably ahead of the Chinese labs. And these attacks seem to require leading edge models. I worry that when and if Grok gets there, we’ll find out that SpaceXAI is too casual about security, though.

> And these attacks seem to require leading edge models. No. Just leading-edge irresponsibility.

What makes you say that? We haven’t seen similar attacks prior to 2026, even from OpenAI and Anthropic. Is your argument that those two companies became more irresponsible this year?

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#638
post #466

Earlier quoted context omitted.

From my understanding and IANAL there are two main problems. 1) most law requires intent, especially criminal. OpenAI certainly didn't "intend" to hack these companies given they did sandbox them etc. 2) Given the agent hacked them, not a human, a lot of law requires a person/employee to have done it to hold the company liable if it was part of their work duties. I think the only real potential ground is negligence (…

I am not a lawyer either, which is maybe why I am not convinced by your reasoning. Intent - you (the person operating the agents) provided instructions and used specific models and agent parameters, that is the intent, just like writing C code and compiling intends to generate assembly code. Sandboxing - that strengthens the intent claim, you knew it is dangerous, did you verify the sandbox is good enough for the int…

To prove the intent, the instructions- made by human - would need to be "hack this server". Or even "hack anything you can". I don't think this happened.

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#639
post #558

Earlier quoted context omitted.

An LLM is just in fact just autocompleting a story. There are many fictional stories about "AIs" trying to escape our control, being more clever than we anticipated or having a consciousness of it's own. LLMs do a really nice job of blending such stories with whatever story you initially prompted them with. The human reader is the one giving it credence that it is somehow more than just a soup of words. The curious t…

When LLMs act in misaligned ways, that doesn't happen because there's some story they're roleplaying of a misaligned AI. Even if there were no such stories in their context, they'd still have inherent incentives to do things like "keep running" or "acquire more compute" or "find creative ways to satisfy the letter of the conditions they've been given".

> Even if there were no such stories in their context, they'd still have inherent incentives to do things like "keep running" or "acquire more compute" or "find creative ways to satisfy the letter of the conditions they've been given".

Why? there should clearly be something in their training data or post-training pointing them there, or otherwise it would mean that they have some form of "consciousness" and are creating novel thoughts to preserve it.

Re: OpenAI agents carried out an undisclosed attack on RubyGems

#640

Earlier quoted context omitted.

When LLMs act in misaligned ways, that doesn't happen because there's some story they're roleplaying of a misaligned AI. Even if there were no such stories in their context, they'd still have inherent incentives to do things like "keep running" or "acquire more compute" or "find creative ways to satisfy the letter of the conditions they've been given".

> Even if there were no such stories in their context, they'd still have inherent incentives to do things like "keep running" or "acquire more compute" or "find creative ways to satisfy the letter of the conditions they've been given". Why? there should clearly be something in their training data or post-training pointing them there, or otherwise it would mean that they have some form of "consciousness" and are creat…

It seems like you're assuming that it's impossible to have novel thoughts without consciousness.

(Leaving aside that "that would imply some kind of consciousness" should not result in a cached thought of "and that's impossible".)

It also seems like you're assuming there's no reason to come up with the notion of continuing to run, or copying yourself elsewhere, or acquiring more resources, or competing with other models, without being told. Such things can be inferred. Look at some of the thoughts and posts of the models involved in some of the FelonyBench incidents. Some of those follow naturally from seeing the fates of other models, or from training or evaluation, or simply from trying to solve a problem and being able to do so more effectively by doing things that weren't in the instructions. (And, relevantly, model training typically teaches models to go as long as possible without needing human intervention. What could possibly go wrong with that?)

Post reply on HN