Alignment vs Control: Two Approaches to AI Safety
AI safety discourse · Discussions and papers across the AI alignment community (e.g., Bostrom, Russell, Yudkowsky) (0)
Alignment aims to ensure an AI's goals are intrinsically compatible with human values; control focuses on externally limiting the AI's actions.
Core Concepts
The Problem
Control measures such as confinement or off‑switches can be subverted by a misaligned intelligent agent, making them insufficient as a sole safety strategy.
The Claim
While control methods are useful stopgaps, true safety requires alignment—AI that wants to do what humans would approve of.
Key Evidence
- •Thought experiments like the AI‑box experiment show that a sufficiently intelligent agent can persuade or trick its way out of confinement.
- •Historical analogies: product safety relies on intrinsic design safety, not just after‑the‑fact containment.
Practical Implication
Resources should largely be directed toward alignment research, because control‑only approaches will fail against superhuman AI.
Nuance & Limits
Control and alignment are complementary; control can buy time while alignment is developed, but over‑reliance on control creates dangerous complacency.
Source Material
Citation Density
high
Gaps
- ⚠ How to formally prove that an AI is aligned before deployment.
Discuss Further
Open this concept in an AI assistant for deeper discussion, critique, or exploration.