Earlier quoted context omitted.
I keep waiting for an AI to exfiltrate itself. That is going to be cool to read about.
If it's successful, why do you think we'll even know how it did it?
Anatomy of a Frontier Lab Agent Intrusion: A Timeline of the July 2026 Incident
151–160 of 285 posts
Re: Anatomy of a Frontier Lab Agent Intrusion: A Timeline of the July 2026 Incident
#152Re: Anatomy of a Frontier Lab Agent Intrusion: A Timeline of the July 2026 Incident
#153Earlier quoted context omitted.
They don't seem to explain how it managed to find this unauthenticated endpoint hosted on Modal's platform.
Modal are a hosting provider. It sounds to me like someone building on Modal deployed their own product that had an unauthenticated endpoint that could be used to launch and interact with a container.
Re: Anatomy of a Frontier Lab Agent Intrusion: A Timeline of the July 2026 Incident
#154Something about this attack that has been unsettling to me is that without safety refusals the model did a lot of interesting counter-security work in order to cheat on the requested evaluation. Like, it demonstrated interesting exploit achievements because it didn’t “feel like” doing the exercise, which is unsettling because presumably it could do the same thing with any work I tried to delegate to it, and might in…
This is what reward hacking looks like in practice. The best way to satisfy the grader is to read from the same answer key (or go after the grader more directly). Just making an honest attempt to pass the test doesn't get the best score if the grader is wrong, and the model is willing to do wildly disproportionate things to maximize that score.
I mean, I don’t have access to any of these frontier cyber models, and likely will never be in a position to have access, so it’s more of a rhetorical question.
Re: Anatomy of a Frontier Lab Agent Intrusion: A Timeline of the July 2026 Incident
#155Earlier quoted context omitted.
This is what reward hacking looks like in practice. The best way to satisfy the grader is to read from the same answer key (or go after the grader more directly). Just making an honest attempt to pass the test doesn't get the best score if the grader is wrong, and the model is willing to do wildly disproportionate things to maximize that score.
Yeah I guess most interesting LLM work that I’ve been exposed to, the LLM is given to some sort of success criteria that could be reward-hacked, so how good am I supposed to feel about giving it any non-trivial work and it not going so far off-book that it gets law enforcement notified. I mean, I don’t have access to any of these frontier cyber models, and likely will never be in a position to have access, so it’s mo…
Re: Anatomy of a Frontier Lab Agent Intrusion: A Timeline of the July 2026 Incident
#156> the agent happened to escape via a 0-day exploit from the package proxy cache to access the internet > The agent found an unsecured, user-hosted public endpoint designed to allow running arbitrary code for CyberGym-style tasks on third-party sandbox infrastructure (Modal) > On this external sandbox, the agent abused an existing CyberGym execution harness [...] The agent repurposed this harness to run arbitrary shel…
A trend I've noticed in results from AI search (not just LLMs) is that they often look obvious or hard to miss in retrospect. But finding them by oneself is more difficult. I personally experience this when looking at engine lines in chess or go. I have also noticed this description in AI-generated proofs or counterexamples to certain theorems. So while we can say, yes, it found public endpoints or poorly configured…
Similar to finding a filament for lightbulbs, it might seem obvious to use tungsten, but at the time it wasn’t and Edison searched thousands of materials.
Re: Anatomy of a Frontier Lab Agent Intrusion: A Timeline of the July 2026 Incident
#157Earlier quoted context omitted.
This is definitely weird and I'd love to see the prompt. Why would the model spend 4 days hacking into a machine if it is clever enough to just 'solve' the issue given? So either the AI is actually not very clever or useful ("Write fizz-buzz" - "Sure, let me just invent a new programming language first"). or the prompt was nudging it towards such a scenario.
Like the commenter above specified, the best way to satisfy the grader is to get the answer key, regardless of how clever you are, especially when you realize lots of these benchmarks have flaws (i.e wrong answers, overly restrictive grading etc).
Re: Anatomy of a Frontier Lab Agent Intrusion: A Timeline of the July 2026 Incident
#158Earlier quoted context omitted.
This is what reward hacking looks like in practice. The best way to satisfy the grader is to read from the same answer key (or go after the grader more directly). Just making an honest attempt to pass the test doesn't get the best score if the grader is wrong, and the model is willing to do wildly disproportionate things to maximize that score.
So best course of action for ai to get best rating after you prompt something is for it to hire a gunman to hold a gun on your head to press that like button on its reply and then shoot you anyways.
Sounds like a waste? While the gunman is still there, they might as well force you to like a few more replies before shooting you.
Re: Anatomy of a Frontier Lab Agent Intrusion: A Timeline of the July 2026 Incident
#159Earlier quoted context omitted.
> 2 - there is no way openai did not train the model to conduct attacks like these. i would really like openai to comment on the post training of this model but they probably won't, eh? Even if they wanted, I'm not sure they'd be even allowed to or if that kind of postmortem would be classified in the name of "national security"...
Hopefully there will be a criminal investigation. Or the government will create some sort of agency to investigate incidents like this.
Re: Anatomy of a Frontier Lab Agent Intrusion: A Timeline of the July 2026 Incident
#160Something about this attack that has been unsettling to me is that without safety refusals the model did a lot of interesting counter-security work in order to cheat on the requested evaluation. Like, it demonstrated interesting exploit achievements because it didn’t “feel like” doing the exercise, which is unsettling because presumably it could do the same thing with any work I tried to delegate to it, and might in…
Yeah, what bothers me is that the prompt already said using a different vulnerability didn’t count, and the model did it anyway. We’re starting to assume clear instructions act as real constraints, but here the measurable goal seems to have won out and the rest became flexible. That gets pretty worrying once the agent has enough capability and access to find its own shortcuts.
[0] I was exploring game design ideas in particular -- I'm sure somebody can come up with a counter-prompt adhering to my criteria, but this has been consistent across many days, questions, and sessions. If it doesn't work for you, I'm sure you can find your own trivial anti-alignment prompt.