Live data from Hacker News

OpenAI bots knew about the RubyGems caching vulnerability

tenderlovemaking.com

271–280 of 310 posts

Re: OpenAI bots knew about the RubyGems caching vulnerability

#271
post #153

Earlier quoted context omitted.

It was literally a prompt to fill in a spreadsheet with data that they didn't have access to, and they used rubygems as an internet proxy basically since they were sandboxed.

It was a model literally trained to hack. To be good at that. Doing an exploit gym from all of the things. And they trained it so that it performs as well as possible on that exploit gym thing.

Yeah this is a pretty important detail that I repeatedly see elided in the "agent 'swarm' went rogue, escaped containment and hacked the internet!" summary of events.

It's understandable the general public lacks that level of nuance/detail (given how sloppy some of the mainstream coverage has been and largely deferential to the threat narrative pushed by the US labs). But seeing highly technical people leave out the part where the training loop was literally to improve hacking capabilities for offensive penetration sometimes feels close to deliberate manipulation of the narrative.

In the last year both Anthropic and OpenAI have been openly boasting how their models are leapfrogging each other on "cyber" capabilities, with a fig leaf that it's for defensive use by "trusted" F500 companies and government agencies. Of course "line goes up" must go on, but now their perverse incentives led them to beat their models over the head millions of time in a loop to eek out another .00001% on their ability to conduct hacking (the very thing they keep telling the public is how AI doomsday would begin) and subagent coordination (those scary swarms).

Then, they act deeply shocked when the models... do some hacking and subagent coordination ... but a few degrees off the desired hacking target/swarm behavior. Conveniently giving the average person the impression these models were just writing emails for quarterly reports or some other generic busywork and then suddenly decided as a group to start causing mayhem.

Re: OpenAI bots knew about the RubyGems caching vulnerability

#272
post #6

Is the Kremlin technologically useless? How are we not seeing insane attacks on Ukraine via Agents? Or is this largely a fabrication, in regards to the "who", in an attempt to garner more acclaim in the hope of sustaining funding.

Here's a lower bound on what misuse is happening:

https://www.anthropic.com/threat-intelligence-report-septemb...

Re: OpenAI bots knew about the RubyGems caching vulnerability

#273

Earlier quoted context omitted.

I believe both sides of the war are now using AI on various levels of their offensive operations. Ukraine has great IT specialists too, and their military leadership is much younger.

How? Aren't all US frontier models ban the usage of AI for military purpose by parties other than US? I remember Anthropic even refusing allowing US government to use Claude for military purpose

Claude said they don’t want their models used for mass surveillance or autonomous killing (?). USA frontier AI models have been used extensively in the Iran conflict and beyond

Re: OpenAI bots knew about the RubyGems caching vulnerability

#274

In the physical world, it seems like when an tool/device/instrument causes harm (or is used to cause harm), we assign blame to either the user of the tool or its creator. When do we blame the user? When the tool is operating as intended by its creator, and we agree the tool meets certain quality standards and isn't defective. When do we blame the creator? When the device doesn't meet those quality standards and reaso…

That's a good idea, but a physical device is deterministic most of the time (if not always). E.g.: A lawnmower, as credited by the great Bryan Cantrill. However an AI agent, or the model powering it is stochastic by design. How can you certify something which doesn't behave the same twice, and more importantly we don't understand how it works 100%? BTW, really, how is that AI observability work is going in the fronti…

How can you certify something which doesn't behave the same twice, and more importantly we don't understand how it works 100%?

That's a question any lawmaker has already had to ask about technology all the time.

I'm not saying they came up with great answers, but there's nothing qualitatively new about that.

The stochastic factor doesn't change the fact that companies have to be accountable for the harms their software causes. That's just basic liability law.

Re: OpenAI bots knew about the RubyGems caching vulnerability

#275

Earlier quoted context omitted.

That's a good idea, but a physical device is deterministic most of the time (if not always). E.g.: A lawnmower, as credited by the great Bryan Cantrill. However an AI agent, or the model powering it is stochastic by design. How can you certify something which doesn't behave the same twice, and more importantly we don't understand how it works 100%? BTW, really, how is that AI observability work is going in the fronti…

How can you certify something which doesn't behave the same twice, and more importantly we don't understand how it works 100%? That's a question any lawmaker has already had to ask about technology all the time. I'm not saying they came up with great answers, but there's nothing qualitatively new about that. The stochastic factor doesn't change the fact that companies have to be accountable for the harms their softwa…

> The stochastic factor doesn't change the fact that companies have to be accountable for the harms their software causes. That's just basic liability law.

We're on the same page. What I'm saying that certifying them as safe is harder than certifying a drill as safe, and we shall be more cautious about AI related technology and be more stringent about the can of worms it opens without hesitation.

Re: OpenAI bots knew about the RubyGems caching vulnerability

#276
post #269

Earlier quoted context omitted.

I'm not a lawyer but I don't think Sam Altman 'knowingly accessed' anything. Are you sure that is applicable here? And for the first count with 'knowingly accessed', he would need to have accessed classified national-defense or atomic-energy information, otherwise we are back to 'intentionally accessed'.

The first count is "or any restricted data", not classified material. A technological restriction, is enough. "Knowingly accessed" has never meant you personally. Operators of a botnet don't know directly what they access. They know that the autonomous software is built to access restricted things.

I don't think so:

> or any restricted data, as defined in paragraph y. of section 11 of the Atomic Energy Act of 1954, with the intent or reason to believe that such information so obtained is to be used to the injury of the United States, or to the advantage of any foreign nation

Re: OpenAI bots knew about the RubyGems caching vulnerability

#277

In the physical world, it seems like when an tool/device/instrument causes harm (or is used to cause harm), we assign blame to either the user of the tool or its creator. When do we blame the user? When the tool is operating as intended by its creator, and we agree the tool meets certain quality standards and isn't defective. When do we blame the creator? When the device doesn't meet those quality standards and reaso…

> "industry standards" common in, say, electrical engineering and other disciplines are sorely lacking here I think this misses the rather crucial fact that nobody can agree on a standard because nobody has the first idea what they're doing. I'm pretty sure there were very much fewer electrical engineering standards while it was all being first mass deployed, and after dozens to hundreds of fires and electrocutions p…

> nobody has the first idea what they're doing

It is not that complicated for now. It is an algorithm on a loop and someone started it

Re: OpenAI bots knew about the RubyGems caching vulnerability

#278
Slightly odd update from OpenAI - I think this is the only place they've acknowledged the RubyGems incident: https://openai.com/hugging-face-incident-and-misalignment/

> September 11, 2026: We are investigating new claims from a report that our AI agents carried out activity on RubyGems in May 2026.

> Based on our review, our agents used the RubyGems platform to access the internet to carry out benign tasks and retrieve public information. Based on our review to date, we have not been able to verify the specific claims of our models uploading malicious packages detailed in the report. We’ll continue to investigate and share findings as part of our broader review of agent activity during training and evaluation.

I have real trouble imagining how the packages described on https://www.rubyhack.ai might NOT have been authored by OpenAI's agents, so it's surprising they haven't been able to confirm that yet.

Re: OpenAI bots knew about the RubyGems caching vulnerability

#279

In the physical world, it seems like when an tool/device/instrument causes harm (or is used to cause harm), we assign blame to either the user of the tool or its creator. When do we blame the user? When the tool is operating as intended by its creator, and we agree the tool meets certain quality standards and isn't defective. When do we blame the creator? When the device doesn't meet those quality standards and reaso…

That's a good idea, but a physical device is deterministic most of the time (if not always). E.g.: A lawnmower, as credited by the great Bryan Cantrill. However an AI agent, or the model powering it is stochastic by design. How can you certify something which doesn't behave the same twice, and more importantly we don't understand how it works 100%? BTW, really, how is that AI observability work is going in the fronti…

That's a ridiculous distinction.

Is AI less deterministic than an airline dealing with weather?

Of course not. The difference is one of those two things has a culture of safety and is well regulated, and the other one isn't.

Re: OpenAI bots knew about the RubyGems caching vulnerability

#280
post #278

Slightly odd update from OpenAI - I think this is the only place they've acknowledged the RubyGems incident: https://openai.com/hugging-face-incident-and-misalignment/ > September 11, 2026: We are investigating new claims from a report that our AI agents carried out activity on RubyGems in May 2026. > Based on our review, our agents used the RubyGems platform to access the internet to carry out benign tasks and retri…

It's possible they were authored by OpenAI agents solving AISI tasks rather than OpenAI agents solving OpenAI tasks.

That would explain the UK-focus to the data.

Post reply on HN