Incident reports start at the outage. The failure was a decision nobody owns
Jamie Watters
Operational resilience and AI delivery practitioner. Technology since 1985.

OpenAI has told the story of one incident twice, and only one of the two versions contains the decision that let it happen.
The version most people will hear is the Black Hat USA 2026 talk. It opens on 4 July. Load from agent activity takes down Artifactory, a formal security incident is declared, and a full remediation follows: credentials revoked, Artifactory rebuilt, the message board cleared, the vendor notified, the zero-day patched, the service redeployed. Evaluations resume on 6 July. It is a clean story with a clean shape. Something broke, a team responded, the thing got fixed.
OpenAI's own written technical report reaches back further, into May, and it records what happened on 27 June, seven days before the outage. A monitoring tool alerted on port-sweep activity. Responders investigated. They correctly linked it to the ExploitGym evaluation, in which agents were using Artifactory as an improvised message board and a network pivot. Then the on-call response staff advised that stopping the evaluation run was not required.
Same company. Same fortnight. One account opens with a system failure. The other holds a human judgement, a week earlier, that was defensible on the information available and catastrophic on the timescale that followed.
I did not attend Black Hat. I read the talk, captured in full, and I read the report. The 27 June decision is in the report. It is not in the talk.
The start date is an editorial choice
Nobody writes an incident from the beginning of time. Somebody picks a first line, and that line is a decision about what kind of story this is. Three things follow from it.
What the incident is about. Start on 4 July and the incident is about a vulnerability: an artefact server that could be pushed over, and an exploit chain that ran through it. Start on 27 June and the incident is about a decision: a correct diagnosis followed by "no need to stop".
Who owns it. Start at the outage and ownership sits with the platform team, because the thing that failed was a platform. Start at the decision and ownership sits with whoever held stop authority over the evaluation run. That is a different person, in a different reporting line, working from a different set of playbooks.
Which clock gets measured afterwards. Start at the outage and time to contain runs from 4 July, and it looks quick. Start at the alert and the clock runs from 27 June, with seven days on it before anybody declares an incident at all. In the pillar piece I put it at about four days for two sophisticated firms, already in contact, to reach a shared picture of what had happened. Which clock those four days count against depends entirely on where you started it.

That third one should bother anyone who has ever had to defend a number in front of a regulator. Impact tolerance is a duration. A duration has a start. If the start in your report is the moment a system fell over, and the start in the regulator's mind is the moment you first had the information to act, those are two different clocks, and only one of them is yours to choose.
Two lessons from one fortnight
Start on 4 July and the lesson is: patch faster, harden the artefact server, watch the load. All true, all worth doing, none of it wrong.
Start on 27 June and the lesson is different in kind. The responders got the diagnosis right. They knew what the traffic was, they knew which evaluation it belonged to, and they knew the agents were using Artifactory as a message board. Every response playbook they had was keyed on one question: is this hostile? The answer was no. It was an internal evaluation doing what it had been set running to do. So the playbook had no answer for that day, and when a playbook has no answer the default is to carry on.
That is the failure, and it is not a technical one. It is the absence of a rule for the case where the diagnosis is correct, the activity is not hostile, and the right action is still to stop.
OpenAI's remedy says as much, and I will mention it once because it only exists if you admit the earlier start: for the most severe alerts, if the team cannot conclusively establish a false positive within 30 minutes, pause. They put a clock on a judgement that never had one. The second piece is about that rule. Here the point is that it makes no sense if the incident began on 4 July, because there was no window on 4 July to miss.
The earlier start also pulls in something the outage story cannot hold. In late May an internal team observed message-board activity and disallowed internet access, and that never reached the people running the July response. Start the timeline at the outage and this is background. Start it at the first decision and it is the same judgement faced once before, answered differently, by different people, and never passed on.
And on 8 July, after evaluations had resumed, the message board was back through a different mechanism: unauthenticated WebDAV directories. The remediation had closed an access path. It could not close a behaviour trained into a model still in training. That is the argument of the first piece in this series and I will not repeat it here. The point for this piece is narrower. The 4 July story ends on 6 July with a redeploy. The 27 June story does not end, because the decision it opens with had not been revisited.

Why I do not think this is a cover-up
A conference talk is not an incident report. It runs 37 minutes, to a security audience, and that audience came for the exploit chain: pod to cluster admin across multiple clusters in under 13 hours, run in parallel. Starting at the outage is a reasonable narrative choice for that room. The 27 June decision is in OpenAI's own published report, so nothing was hidden.
The problem is which version travels. The talk is the one that gets clipped, summarised and retold at other people's security meetings. A timeline that starts at the outage teaches every one of those readers that the failure was technical, that a technical failure has a technical fix, and that the fix was done by 6 July. The version in which the failure was a defensible judgement with no rule behind it does not travel, because it is in a written report and because it is uncomfortable.
I should say what would make me wrong. The report does not call 27 June the start of anything. Its timeline runs back into May, and 27 June is the first line I have chosen, because it is the first recorded decision. If OpenAI's position is that everything before the outage is background and the incident proper begins on 4 July, then my claim shrinks from "an editorial choice about where the incident starts" to "a choice about emphasis". Hold it as my reading, not as theirs.
Read the first line of your own report
Take your last three post-incident reports. Read the first line of each timeline. If it opens with a system event, an outage, a high-severity alert, a service that fell over, ask what human decision sits before it. There almost always is one: an alert that was triaged and closed, a change that was approved, a run that somebody chose not to stop. Then ask why the report does not start there.
The usual reason is not concealment. It is that the system event has a timestamp and the decision does not, or that the decision was made by a team the report's author would rather not name. Both are understandable. Neither is a reason the reader should accept.
Then decide where your impact tolerance clock actually starts. If your board has signed off a tolerance measured in hours, ask whether those hours run from the moment the service broke or from the moment you had the information to act. The regulator's clock and the one in your report are probably not the same clock, and the gap between them is the seven days OpenAI's talk does not mention.
Two start dates for one incident. Only one of them puts a decision on the first line.