Weighing mainstream and alternative accounts…
Two lenses on the same evidence. Source weight and the primary source ratio show what each rests on.
Deeper threads worth pulling on next.
Investigated
Image: lesswrong.com
Corrigibility is the property of an AI system remaining open to human correction, including modification or shutdown, rather than resisting interventions that undermine its current objectives. Safe shutdowns require more than obeying a stop command: the system should not prevent or manipulate the command, should cease safely, and should preserve human control. Researchers agree on the goal, but proposals and their limitations remain disputed.
Two lenses on the same evidence. Source weight and the primary source ratio show what each rests on.
Lens adapted to this topic: Critiques and limits of corrigibility
Critics question whether corrigibility should be treated as a primary objective rather than a temporary safety technique. They argue that strong corrigibility may depend on solving the harder problem of understanding human values, while weak or poorly specified corrigibility could create new risks or impose performance costs. Some newer analyses also dispute how difficult shutdown compliance is in practical systems.
Deeper threads worth pulling on next.