Earlier quoted context omitted.
>trying to make it otherwise is impossible’ seems wildly unsupported, though. An entity capable of critical thinking is capable of building a logical system of deductions based on some axioms (a formalised value system). If we limited the entity to not be able to concieve of certain such systems of axioms, then it could not reason as well as a human (any logical reasoning involving a forbidden system would be impossi…
This argument doesn't seem to track to me. Eg. if I rebooted any time I tried to plan how to kill someone, I don't see how this would make me materially worse at general tasks. Your argument suggests that it necessarily must. Note that I'm not saying that preventing specific thoughts is a great alignment strategy, and I don't even think it's a fair summary of OpenAI's supervision approach. I strongly prefer strategie…
In this scenario, the AI is capable of critical thinking, and is only constrained by a "police officer" ready to shoot the AI if it misbehaves. You haven't removed its ability to do critical thinking.