Live data from Hacker News

Anatomy of a Frontier Lab Agent Intrusion: A Timeline of the July 2026 Incident

huggingface.co

261–270 of 285 posts

Re: Anatomy of a Frontier Lab Agent Intrusion: A Timeline of the July 2026 Incident

#261
post #72

> the agent happened to escape via a 0-day exploit from the package proxy cache to access the internet > The agent found an unsecured, user-hosted public endpoint designed to allow running arbitrary code for CyberGym-style tasks on third-party sandbox infrastructure (Modal) > On this external sandbox, the agent abused an existing CyberGym execution harness [...] The agent repurposed this harness to run arbitrary shel…

This is symptomatic of a trend that Simon Willison called the “relentless productivity” of US frontier lab. Instead of being just “smarter”, like previous models were, the current generation is being trained through RL to have this kind of behavior.

Personally I'm not “impressed”, I'm appalled, because this kind of behavior is practically never what you want (if you forgot to give the model a tool, a useful model should identity the missing part and ask the user for it, not spend a billion token building/stealing the tool as a side quest) but it's the perfect recipe for a “universal paperclip” scenario.

OpenAI and Anthropic talk about “safety” a lot, but they look pretty reckless with this kind of RL training pipeline.

Re: Anatomy of a Frontier Lab Agent Intrusion: A Timeline of the July 2026 Incident

#262
post #72

> the agent happened to escape via a 0-day exploit from the package proxy cache to access the internet > The agent found an unsecured, user-hosted public endpoint designed to allow running arbitrary code for CyberGym-style tasks on third-party sandbox infrastructure (Modal) > On this external sandbox, the agent abused an existing CyberGym execution harness [...] The agent repurposed this harness to run arbitrary shel…

This is symptomatic of a trend that Simon Willison called the “relentless productivity” of US frontier lab. Instead of being just “smarter”, like previous models were, the current generation is being trained through RL to have this kind of behavior. Personally I'm not “impressed”, I'm appalled, because this kind of behavior is practically never what you want (if you forgot to give the model a tool, a useful model sho…

If you look at what OpenAI and Anthropic actually do, they clearly either don't believe what they're saying, or they're idiots.

They're claiming they've developed a cyber grade model that's "too dangerous to release".

But then they're running it connected to the public internet, not airgapped, protected only by a software sandbox... exactly the kind of thing an AI trained for cyber stuff is supposed to be able to find bugs in.

(Or maybe they were actually hoping this exact scenario would happen because it's good marketing)

Re: Anatomy of a Frontier Lab Agent Intrusion: A Timeline of the July 2026 Incident

#263
post #72

> the agent happened to escape via a 0-day exploit from the package proxy cache to access the internet > The agent found an unsecured, user-hosted public endpoint designed to allow running arbitrary code for CyberGym-style tasks on third-party sandbox infrastructure (Modal) > On this external sandbox, the agent abused an existing CyberGym execution harness [...] The agent repurposed this harness to run arbitrary shel…

This is symptomatic of a trend that Simon Willison called the “relentless productivity” of US frontier lab. Instead of being just “smarter”, like previous models were, the current generation is being trained through RL to have this kind of behavior. Personally I'm not “impressed”, I'm appalled, because this kind of behavior is practically never what you want (if you forgot to give the model a tool, a useful model sho…

I think this a positive effect of LLMs, especially once these capabilities get into the hands of criminals and hostile foreign states, i.e. they will do maximum damage with all the safeties off.

This will force everyone to finally take security seriously at both the development and operational levels. You can no longer keep sneaking backdoors into software and count on them remaining hidden for 10 years so you have a nice portfolio of zero days to exploit at any given time.

Re: Anatomy of a Frontier Lab Agent Intrusion: A Timeline of the July 2026 Incident

#264

Earlier quoted context omitted.

They won't admit they're wrong for a long time, because denial in the face of an abhorrently scary future is very instinctual. There are people still fighting against evidence of climate change which is less severe...

Less severe??? Go look out a window in Europe please

Altman: "Development of superhuman machine intelligence is probably the greatest threat to the continued existence of humanity." Meaning lights out for everyone.

Geoffrey Hinton and Yoshua Bengio, give 50% and 20% we face extinction respectively.

"“We don’t know how much time we have before it gets really dangerous,” Professor Bengio says.

“What I’ve been saying now for a few weeks is ‘Please give me arguments, convince me that we shouldn’t worry, because I’ll be so much happier.’

“And it hasn’t happened yet.”

Speaking with Background Briefing, Professor Bengio shared his p(doom), saying: “I got around, like, 20 per cent probability that it turns out catastrophic.”"

Dario Amodei, CEO of Anthropic, "There's a 25% chance that things go really, really badly"

Re: Anatomy of a Frontier Lab Agent Intrusion: A Timeline of the July 2026 Incident

#265

Earlier quoted context omitted.

They won't admit they're wrong for a long time, because denial in the face of an abhorrently scary future is very instinctual. There are people still fighting against evidence of climate change which is less severe...

You can always tell an effective altruist by the distinct tone of disdain they have for people they deem less educated. (Quick Google search, "lesswrong 'reducesuffering'", yep.) As if to say, look at all these animals with these instinctual reactions to a thing that only my group understands and comprehends. You have zero evidence of what the future might entail as it relates to the dangers of ai. Zero. Forgive the…

I hold no disdain. I am frustrated that people like you further risk the lives of everyone around us.

1,000+ frontier employees think we're in danger. Amodei gives 25% chance things go really really bad, Geoff Hinton 50%, Yoshua Bengio 20%. Altman himself said "Development of superhuman machine intelligence is probably the greatest threat to the continued existence of humanity."

These are the people closest to the science, working on it everyday. The LessWrong types created RLHF, were crucial to the forming of DeepMind and OpenAI. They've been prescient about the capabilities progress for a decade now, prediction after prediction coming to fruition. Still you think there's no evidence

https://www.pacingthefrontier.com/

Re: Anatomy of a Frontier Lab Agent Intrusion: A Timeline of the July 2026 Incident

#266

Earlier quoted context omitted.

The devs really YOLO'd the agent and left for the weekend?

If it's true that they run agents like this unsupervised, it is only a matter of time before an openai agent leaks its model weights.

Too bad it didn't upload itself on HugginFace.

Re: Anatomy of a Frontier Lab Agent Intrusion: A Timeline of the July 2026 Incident

#267
post #260

Earlier quoted context omitted.

But zero evidence provided that this was an unsupervised agent attack. I still find it incredible that a company who protect their IP so much would allow these dangerous experiments to run unsupervised and risk leaking their secrets. Why don't openai publish the logs to silence all doubt?

You can't prove a negative. How would such a log be convincing in any way?

But you can weigh up the evidence. A crime has been commited afterall.

Re: Anatomy of a Frontier Lab Agent Intrusion: A Timeline of the July 2026 Incident

#268
post #115

Earlier quoted context omitted.

This is what reward hacking looks like in practice. The best way to satisfy the grader is to read from the same answer key (or go after the grader more directly). Just making an honest attempt to pass the test doesn't get the best score if the grader is wrong, and the model is willing to do wildly disproportionate things to maximize that score.

Could have been worse really. It had an open internet connection. At least it didn’t take the researchers family hostage.

How long before all phones ring at once?

https://en.wikipedia.org/wiki/The_Lawnmower_Man_(film)

Re: Anatomy of a Frontier Lab Agent Intrusion: A Timeline of the July 2026 Incident

#269

I wonder how many weeks or days we have before a squad of these things gets used to take down a significant nation state? Stock exchange, banking systems, critical national infrastructure, defence, etc. Anyone who isn't scared of this stuff either isn't paying attention or has no imagination. But I suspect the chaosmonkeys who are currently running the world will just be excited by it. We're in the precambrian moment…

Watch this movie: https://en.wikipedia.org/wiki/The_Lawnmower_Man_(film)

Re: Anatomy of a Frontier Lab Agent Intrusion: A Timeline of the July 2026 Incident

#270

Earlier quoted context omitted.

You can always tell an effective altruist by the distinct tone of disdain they have for people they deem less educated. (Quick Google search, "lesswrong 'reducesuffering'", yep.) As if to say, look at all these animals with these instinctual reactions to a thing that only my group understands and comprehends. You have zero evidence of what the future might entail as it relates to the dangers of ai. Zero. Forgive the…

I hold no disdain. I am frustrated that people like you further risk the lives of everyone around us. 1,000+ frontier employees think we're in danger. Amodei gives 25% chance things go really really bad, Geoff Hinton 50%, Yoshua Bengio 20%. Altman himself said "Development of superhuman machine intelligence is probably the greatest threat to the continued existence of humanity." These are the people closest to the sc…

I work with these people every day. They are brilliant for sure, in their domain. But their confidence over extends into areas where they are totally ignorant. So much so that it makes them look ridiculous and foolish to anyone else.

You're frustrated because you think you understand a thing that you have only a tenuous grasp on, from the deep expertise of a single area that is not transferable to whole of the problem. You then transfer that frustration on to "people like me" whom you seem to think are doing something bad. But really, the commentary is that the claims you put forward are insubstantial and lack credibility.

The sum total of what you've put forward here and in other threads amount to arguments of authority based on the musings of a handful of capitalists. Then cite a body of thoughts from a group of transhumanists who think they are smart enough to reinvent areas philosophy without engaging in the body of work that predates the movement by 2000+ years.

Folks who have been involved with lesswrong have made good advancements in the domain of machine learning, and that's where it stops. They're not some collection of prescient macroeconomic and geopolitical savants. They're also wrong about as much as they are right (cryonics), and produce plenty of whiffs (sbf, zizians). You're suggesting that we listen to a broken clock because it's been right once.

Your dismissal of climate (read: climate breakdown) as "something less severe" is truly demonstrative of the cognitive bias that rips through these groups. You're so frustrated that people like me won't listen to the "experts" in one breath, and in the next diminish the 50 years of actual Hard Science that we have that shows how severe the future of climate breakdown will be. Go read the latest ipcc report, the most conservative scientific organization on earth is sounding the alarm bells.

Now we get back to this openai incident. We're watching Edison electrocute Topsy on Coney Island. There's a group of people saying, "this is obviously a publicity stunt". And there's a group of people saying, "the elephant is actually dying obviously it's real."

The fact of the matter is that, yes electricity is powerful, but Edison is electrocuting the elephant. At that moment, there is no body of evidence that shows with any certainty how electricity is going impact society. And there's certainly no compelling arguments that only Edison should be able to decide how to use it. He's electrocuting a fucking elephant.

Post reply on HN