Earlier quoted context omitted.
>Now, tbh it does rely on the assumption that we are actually seeing more undesired behaviour than before. If that was false then it would falsify the Waluigi hypothesis. This is exactly my point. There is no evidence given that we are seeing more Waluiginess post-RLHF than we did pre-RLHF. The competing hypothesis seeks to explain the behavior we actually have evidence for, which is "it is disappointingly easy to el…
> There is no evidence given that we are seeing more Waluiginess post-RLHF than we did pre-RLHF. Testing with the non-RLHF GPT 3.5 API you could probably figure out whether there's more or less Waluiginess, but you're right they post doesn't present this.
There is no such API, though, is there? AFAIK, GPT-3.5-turbo, either the updated or snapshot version, is the RLHF model (but bring your own “system prompt”.)