Earlier quoted context omitted.
We can seperate them but the $ value of an agent that does is much lower than one that doesn't. As a pre LLM analogy imagine working at a bank with a whitelist firewall. You need to install a package but requires an IT ticket. Safer but slooooower. Now not saying what the answer here is but that is the issue. The answer may be more like industries that get safer through lessons (like aviation) rather than go for 100%…
what? Aviation safety is not designed to get safer through lessons? They literally try to ensure it is 100% safe out of the gate. The accidents that happen are usually statistical outliers and lead to loss of life. That's what it means when they say aviation regulations are written in blood. Not that they just fling planes into the sky and be like "boy i hope we learn some new regulations from this". The number of ai…
Lockdown Mode
11–20 of 39 posts
Re: Lockdown Mode
#12So we still don't have a reliable way to separate instructions from data when talking to an LLM, a problem that humans learned how to solve decades ago in areas like SQL and memory safety. But hey, we have these hopefully-not-leaky containers, which are probably implemented with just more system prompts. How long until somebody figures out how to trick Codex into disabling Lockdown Mode for you?
Humans also do not know how to do this reliably, which is why phishing is still a thing and always will be.
Re: Lockdown Mode
#13The existence of lockdown mode does however imply that ChatGPT, in its default settings, does not provide robust protection against sufficiently determined data exfiltration attacks!
Re: Lockdown Mode
#14Re: Lockdown Mode
#15So we still don't have a reliable way to separate instructions from data when talking to an LLM, a problem that humans learned how to solve decades ago in areas like SQL and memory safety. But hey, we have these hopefully-not-leaky containers, which are probably implemented with just more system prompts. How long until somebody figures out how to trick Codex into disabling Lockdown Mode for you?
> So we still don't have a reliable way to separate instructions from data when talking to an LLM Humans also do not know how to do this reliably, which is why phishing is still a thing and always will be.
Re: Lockdown Mode
#16On the one hand this is exactly the right solution to prevent lethal trifecta exfiltration attacks. The existence of lockdown mode does however imply that ChatGPT, in its default settings, does not provide robust protection against sufficiently determined data exfiltration attacks!
Re: Lockdown Mode
#17I have mixed feelings about this feature. We're playing with tech that's supposed to do human-shaped things but can't be trusted nearly as much as a human employee (and can't be held responsible for what it does). Restricting the tools available to that patently untrustworthy entity doesn't solve the problem, it just makes the entity less useful, forcing you to sooner or later let it out of the jail.
Re: Lockdown Mode
#18On the one hand this is exactly the right solution to prevent lethal trifecta exfiltration attacks. The existence of lockdown mode does however imply that ChatGPT, in its default settings, does not provide robust protection against sufficiently determined data exfiltration attacks!
Related: Simon Willison’s post on OpenAI’s new Lockdown Mode (he coined the “lethal trifecta” term this is based on): https://simonwillison.net/2026/Jun/5/openai-help-lockdown-mo...
Re: Lockdown Mode
#19https://x.com/sama/status/1891533802779910471
Re: Lockdown Mode
#20Earlier quoted context omitted.
Related: Simon Willison’s post on OpenAI’s new Lockdown Mode (he coined the “lethal trifecta” term this is based on): https://simonwillison.net/2026/Jun/5/openai-help-lockdown-mo...
Related: simonw is Simon Willison