Earlier quoted context omitted.
We don’t expect 100% reliability from humans-humans will slack off, steal, defraud, harass each other, sell your source code to a foreign intelligence service, turn your business behind your back into a front for international drug cartels-some of that is very low probability, but never zero probability-so is it really a problem if we can’t reduce the probability to literally zero for AIs either?
Humans have incentives to not do those things. Family. Jail. Money. Food. Bonuses. Etc. If we could align an AI with incentives in the same way we can a person then youd have a point. So far alignment research is hitting dead ends no matter what fake incentives we try to feed an AI.
Re: AI documentation you can talk to, for every repo
#131Can you remind me of the link between alignment and writing accurate documentation? Honestly don't understand how they are linked.