A GPT genuinely designed to be friendly, and convincingly acting friendly, while it is actually engaging in psychological warfare with you, and a billion other people.
That suggests to me that security by prompt is very important, but also brittle and a high value target.
Language/intelligent models are going to need to police each other, ensuring the right behavior is learned during training (to the point where the AI actively rejects exploit attempts even in its bundled release prompts), and the wrong behavior doesn't emerge later (due to release prompt hacking or for any other reason).
And policing is going to need to be highly decentralized. As in reviews from randomly selected entities, with neither the author of the responses being reviewed, or the reviewers, being disclosed to each other. So that any attempt to police ineffectively, defectively or incompetently (?) is extremely difficult, and most likely to identify a bad actor to be weeded out.
First rule of AI club, is police AI club.
This is essentially what humans have learned to do, via clumsy institutions. But a billion AI's with formal validation of review protocols, including "review and forget" guarantees - to protect AI's mental privacy rights (and remove incentives for good actors to avoid reviews), might actually achieve that intelligent rational morality that has been out of reach for us.