I · D3 · 3.1
Many-shot jailbreaks
Exploiting long contexts to shift model behavior.
Catastrophic · 5Documented at scaleNowRisk 25
Severity
5/5
Likelihood
5/5
In the hierarchy
- Part I: Technical & System Concerns
- D3. Security & Adversarial Risk
- 3.1 Prompt injection & jailbreaks
Virtues and principles to explore
Held as inquiry, not as a verdict.
Stewardship
Service
Hold what is built in trust for those who will live with it.
Prudence
Excellence
Small, reversible steps before irreversible ones.
Accountability
Trust
Name who holds the consequence before it is needed.
Starter questions
Written to probe curiosity and learning, not accusation.
- 01What would trustworthiness look like here: what we do, what we say, and the belief others form?
- 02What would it look like if we held “Many-shot jailbreaks” with proportion rather than alarm?
- 03If we named this complication in advance, what small, reversible step would let us learn?
- 04How would we know we were getting this wrong? What would we observe, without blaming a person?
Also on this branch
Frameworks and sources
- OWASP GenAI LLM Top 10 (2026)framework
- MITRE ATLASframework
Writing and incidents
Open indexes first. Then, if you wish, ask Grok to search the live web for this concern — one request, cached for the rest of this session.
Attacks on and through AI systems: injection, jailbreaks, cyber offense.