Come down a rabbit hole with OpenAI o1 system card
1–2 of 2 posts
Re: Come down a rabbit hole with OpenAI o1 system card
#2The most interesting findings (at least from my perspective)
- OpenAI measured "honesty" that could become an interesting model eval others should adopt.
- Using instruction heirarchy to _avoid_ prompt injections
- When pushed to pursue misaligned goals, it tried to disable oversight in 5% of tests!