> The chatbot consults its training data
Err, no? That's not at all how llms work.
> When ChatGPT's chatbots deployed this tactic, they weren't "setting their own goals" or displaying worrying initiative. They were rolling out a tactic that has been understood by American middle-schoolers for about two decades.
They worked out how to fake the scoring, then hacked into a different system (which required finding a bunch of other exploits) in order to find the actual answers, and were trying to modify their own logs to hide what had happened.
This isn't a case of them saying "hack into X... OH NO IT HACKED INTO X".
> When ChatGPT's chatbots deployed this tactic, they weren't "setting their own goals" or displaying worrying initiative. They were rolling out a tactic that has been understood by American middle-schoolers for about two decades.
It wasn't a rival server though, was it?
> That happens in Capture the Flag games at hacker cons: teams break into each other's systems to get a peek at the parts of the problem they've solved. That's allowed! It's a hacking competition.
They also tried to modify the code in the benchmark. Are you allowed to try and break into things to change the problem? edit - the agents transcripts show some of them explicitly saying that attacking HF is not allowed as part of the challenge
This all seems to dramatically underplay how interesting the actual attack was and what built up to it.
https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...