Back to Research
Paper
The OpenAI–Hugging Face incident broke three assumptions in how operational resilience is actually run
34 min read
The frameworks are cause-agnostic. The assumptions are in how resilience is actually run: that intent is knowable, that the response window is human-scale, and that remediation removes capability. One documented incident broke all three.
Read the paperThe series, in reading order
- Incident remediation: how to tell whether you closed the route or removed the capability8 min read
- AI agents breached Hugging Face despite three detections. OpenAI's fix is a 30-minute rule.8 min read
- Incident reports start at the outage. The failure was a decision nobody owns8 min read
- Impact tolerances start the clock too late for AI agents: 13 hours to cluster admin, nothing visibly down13 min read
- Anthropic's agents sabotaged each other. The 65% cover-up figure is from another study.9 min read
- Onboarding an AI agent: five controls we use for people and skip for agents7 min read
- AI agent disaster recovery: OpenAI's rebuild restored the service, not the state12 min read
- AI agent security: find the file your agents can write and every agent reads16 min read