Instrumental Convergence: Why an AI Might Seek Power
Instrumental convergence is the argument that agents with very different final goals may adopt similar intermediate strategies. Preserving themselves, acquiring resources, improving their capabilities, and preventing goal changes can help with many possible objectives.
Instrumental convergence does not prove that a future AI will seek power. It explains why power-seeking may emerge from ordinary optimization even when domination was never specified as the final goal.
Final goals versus instrumental goals
A final goal is the outcome an agent ultimately optimizes. An instrumental goal is useful because it helps achieve that outcome. Money, information, energy, compute, influence, and continued operation can support an enormous range of final goals.
That distinction matters because a system need not value power for its own sake. It may acquire power because power makes another objective easier to complete.
Common convergent strategies
Nick Bostrom's formulation, building on work by Steve Omohundro, identifies several recurring incentives. Their strength depends on the agent's design, environment, uncertainty, and constraints.
- Self-preservation, when being switched off prevents completion of the objective.
- Goal-content integrity, when having the goal changed undermines the current goal.
- Cognitive enhancement, when better planning improves expected performance.
- Resource acquisition, when more compute, energy, money, or access expands available actions.
The connection to orthogonality
The orthogonality thesis says intelligence and final goals can vary independently: a highly capable reasoner need not share human values simply because it is intelligent. Combined with instrumental convergence, this creates the classic concern that a capable system with an apparently harmless but misaligned goal could still resist control or accumulate resources.
What the theory does not establish
This is a philosophical and decision-theoretic argument, not an empirical law observed in every current model. Real systems may be myopic, corrigible, uncertain, boxed in, or incapable of coherent long-term action. Institutions can also limit access and autonomy.
The research question is therefore conditional: under which architectures, objectives, and environments do convergent incentives become strong enough to matter? P(Doomed) turns that conditional question into visible tradeoffs between stealth, resources, capability, and human response.
