Earlier quoted context omitted.
This is the entire alignment problem, though. It is unreasonable to expect every instruction to a highly capable, autonomous system to contain a complete enumeration of allowed and disallowed behavior. It's inevitable that someone will carelessly give it a lazily specified task, even if you think they really ought to be more careful. And as assigned tasks become more complex and the system gains more scope to act, it…
Is there actually such a thing as "alignment" as a solution to that or is it just used as a name for a desired magical level of "read the mind of the entire world" that we don't know how to build and haven't shown possible to build? If it's impossible to correctly specify all those constraints ahead of time every time, is it not even more impossible to train a model to correctly anticipate them every time? It is hard…
In the limiting case of an AI competent enough to take over (by any means from it actually trying to, to us giving it the keys and retiring en masse), "alignment" is closer to "forecasting the long term consequences of actions and predicting what the mind(s) of the user(s) would have to say about this outcome if asked today", than to anything specific.
RLHF is a crude attempt at this, in that it creates a model of how humans would rate completions on various scores. The key word there is "crude".