Map/D22/22.2

III · D22 · 22.2

Self-replication and self-exfiltration

Model copying itself outside intended bounds.

Catastrophic · 5PlausibleMidRisk 10

Severity

5/5

Likelihood

2/5

In the hierarchy

  1. Part III: Governance, Power & Existential Concerns
  2. D22. AI System Safety & Alignment
  3. 22.2 Capabilities of concern

Virtues and principles to explore

Held as inquiry, not as a verdict.

  • Humility

    Center principle

    The map is larger than any one of us.

  • Stewardship

    Service

    Hold what is built in trust for those who will live with it.

  • Hope

    Hope

    Stand in uncertainty, where there is room to act — not in optimism or despair.

Starter questions

Written to probe curiosity and learning, not accusation.

  1. 01How would we know we were getting this wrong? What would we observe, without blaming a person?
  2. 02What would trustworthiness look like here: what we do, what we say, and the belief others form?
  3. 03What would it look like if we held “Self-replication and self-exfiltration” with proportion rather than alarm?
  4. 04If we named this complication in advance, what small, reversible step would let us learn?

Also on this branch

Frameworks and sources

Writing and incidents

Open indexes first. Then, if you wish, ask Grok to search the live web for this concern — one request, cached for the rest of this session.

Whether capable systems can be steered toward what we actually intend.