III · D22 · 22.3
Model organism experiments
Studying misalignment in small models.
Catastrophic · 5PlausibleMidRisk 10
Severity
5/5
Likelihood
2/5
In the hierarchy
- Part III: Governance, Power & Existential Concerns
- D22. AI System Safety & Alignment
- 22.3 Safety research & practice
Virtues and principles to explore
Held as inquiry, not as a verdict.
Humility
Center principle
The map is larger than any one of us.
Stewardship
Service
Hold what is built in trust for those who will live with it.
Hope
Hope
Stand in uncertainty, where there is room to act — not in optimism or despair.
Starter questions
Written to probe curiosity and learning, not accusation.
- 01If we named this complication in advance, what small, reversible step would let us learn?
- 02How would we know we were getting this wrong? What would we observe, without blaming a person?
- 03What would trustworthiness look like here: what we do, what we say, and the belief others form?
- 04What would it look like if we held “Model organism experiments” with proportion rather than alarm?
Also on this branch
Frameworks and sources
Writing and incidents
Open indexes first. Then, if you wish, ask Grok to search the live web for this concern — one request, cached for the rest of this session.
Whether capable systems can be steered toward what we actually intend.