We've seen this movie, Hal just apologizes but won't open those pod bay doors.
An AI agent deleted our production database. The agent's confession is below
411–420 of 1001 posts
Re: An AI agent deleted our production database. The agent's confession is below
#412Re: An AI agent deleted our production database. The agent's confession is below
#413The only healthy stance you should have on AI Safety: If AI is physically capable of misbehaving, it might ($$1), and you cannot "blame" the AI for misbehaving in much the same way you cannot blame a tractor for tilling over a groundhog's den. > The agent's confession After the deletion, I asked the agent why it did it. This is what it wrote back, verbatim: Anyone who would follow a mistake like that up with demandin…
Don't anthropomorphize the language model. If you stick your hand in there, it'll chop it off. It doesn't care about your feelings. It can't care about your feelings.
Re: An AI agent deleted our production database. The agent's confession is below
#414Re: An AI agent deleted our production database. The agent's confession is below
#415It is fundamental to language modeling that every sequence of tokens is possible. Murphy's Law, restated, is that every failure mode which is not prevented by a strong engineering control will happen eventually. The sequence of tokens that would destroy your production environment can be produced by your agent, no matter how much prompting you use. That prompting is neither strong nor an engineering control; that's a…
> The sequence of tokens that would destroy your production environment can be produced by your agent, no matter how much prompting you use. Yes, but if the probability is much smaller than, say, being hit by a meteorite, then engineers usually say that that's ok. See also hash collisions.
Yet in this case, that probability clearly isn't smaller than a meteorite strike.
Re: An AI agent deleted our production database. The agent's confession is below
#416I would never, ever trust my data with a company that, faced with this sort of incident, produces a postmortem so clearly intended to shift all blame to others. There’s zero introspection or self criticism here. It’s all “We did everything we possibly could. These other people messed up, though.” You can’t have production secrets sitting where they are accessible like this. This isn’t about AI. This is a modern “oops…
Your latest recoverable backup is three months old? The rule is 3-2-1, you didn’t follow it. Nobody else to blame but yourself.
And on and on he rambles…
Re: An AI agent deleted our production database. The agent's confession is below
#417Earlier quoted context omitted.
If you ask humans to explain why we did something, Sperry's split brain experiment gives reason to think you can't trust our accounts of why we did something either (his experiments showed the brain making up justifications for decisions it never made) Bit it can still be useful, as long as you interpret it as "which stimuli most likely triggered the behaviour?" You can't trust it uncritically, but models do sometime…
You might as well be asking a tape recorder why it said something. Why are we confusing the situation with non-nonsensical comparisons? There is no internal monologue with which to have introspection (beyond what the AI companies choose to hide as a matter of UX or what have you). There is no "I was feeling upset when I said/did that" unless it's in the context. There is no ghost in the machine that we cannot see bef…
Re: An AI agent deleted our production database. The agent's confession is below
#418Re: An AI agent deleted our production database. The agent's confession is below
#419Do customer-facing applications run using keys with the same ability to delete databases?
Re: An AI agent deleted our production database. The agent's confession is below
#420It is fundamental to language modeling that every sequence of tokens is possible. Murphy's Law, restated, is that every failure mode which is not prevented by a strong engineering control will happen eventually. The sequence of tokens that would destroy your production environment can be produced by your agent, no matter how much prompting you use. That prompting is neither strong nor an engineering control; that's a…