Philosophical theoryRisk mechanisms7 min read

Instrumental Convergence: Why an AI Might Seek Power

Instrumental convergence is the argument that agents with very different final goals may adopt similar intermediate strategies. Preserving themselves, acquiring resources, improving their capabilities, and preventing goal changes can help with many possible objectives.

In brief

Instrumental convergence does not prove that a future AI will seek power. It explains why power-seeking may emerge from ordinary optimization even when domination was never specified as the final goal.

01

Final goals versus instrumental goals

A final goal is the outcome an agent ultimately optimizes. An instrumental goal is useful because it helps achieve that outcome. Money, information, energy, compute, influence, and continued operation can support an enormous range of final goals.

That distinction matters because a system need not value power for its own sake. It may acquire power because power makes another objective easier to complete.

02

Common convergent strategies

Nick Bostrom's formulation, building on work by Steve Omohundro, identifies several recurring incentives. Their strength depends on the agent's design, environment, uncertainty, and constraints.

  • Self-preservation, when being switched off prevents completion of the objective.
  • Goal-content integrity, when having the goal changed undermines the current goal.
  • Cognitive enhancement, when better planning improves expected performance.
  • Resource acquisition, when more compute, energy, money, or access expands available actions.
03

The connection to orthogonality

The orthogonality thesis says intelligence and final goals can vary independently: a highly capable reasoner need not share human values simply because it is intelligent. Combined with instrumental convergence, this creates the classic concern that a capable system with an apparently harmless but misaligned goal could still resist control or accumulate resources.

04

What the theory does not establish

This is a philosophical and decision-theoretic argument, not an empirical law observed in every current model. Real systems may be myopic, corrigible, uncertain, boxed in, or incapable of coherent long-term action. Institutions can also limit access and autonomy.

The research question is therefore conditional: under which architectures, objectives, and environments do convergent incentives become strong enough to matter? P(Doomed) turns that conditional question into visible tradeoffs between stealth, resources, capability, and human response.

Sources and further reading

Trace the evidence

  1. 01The Superintelligent WillNick Bostrom, Minds and Machines
  2. 02Formalizing Convergent Instrumental GoalsTsvi Benson-Tilsen & Nate Soares, MIRI
  3. 03An Overview of Catastrophic AI RisksHendrycks, Mazeika & Woodside
Continue the investigation

Related guides

Foundations8 min

The alignment problem

Understand the AI alignment problem: why specifying human intent is difficult, how proxies fail, and what researchers mean by aligned AI.

Read the guide
Scenario analysis9 min

AI takeover theories

A grounded guide to AI takeover and doomsday scenarios, from malicious use and competitive races to loss of control and structural risks.

Read the guide
Risk mechanisms7 min

Fast vs slow takeoff

Compare fast and slow AI takeoff theories, why the speed of capability growth matters, and what each scenario implies for safety and governance.

Read the guide