Ray Poynter, 4 September, 2026


Several leading AI services ran into trouble within the same few hours on Thursday this week. OpenAI’s status page recorded elevated errors across ChatGPT and Codex for roughly two hours before the incident was resolved. Anthropic separately reported errors affecting several Claude models, with the main incident closing after nearly three hours. Grok and other services were reported as unavailable over a similar period.

The timing invited speculation about a shared cause. The public evidence supports something more modest. OpenAI and Anthropic pointed to different things in their status reports, and coincidence, congestion and separate faults all remain possible/probable explanations.

The primary sources are worth a look if you want to see how the two companies describe these events: OpenAI’s incident record and Anthropic’s status history.

Why this matters more than it did a year ago

Three trends ask a resilience question.

Models are more capable, so each one now carries a larger share of the workflow. When a tool only drafted a summary, an outage cost you an hour. When it creates your analysis, cleans your open ends and drafts your client deliverables, an outage stops the project.

Infrastructure is concentrating. Acquisitions and platform partnerships mean that apparently separate services can rest on shared compute, shared networks and shared providers. Diversity at the brand level can conceal similarity underneath.

Attack capability is rising too. Agents that can accelerate legitimate work can also accelerate hostile work, which raises the odds that a future interruption is deliberate rather than accidental.

Put those together and the cost of one service becoming unavailable, or becoming untrustworthy, keeps climbing.

What to do about it

Here are some key steps you can take.

  1. Map the dependency. List the research processes that now rely on an AI service. Most teams find more of them than they expected, because adoption happens person by person.
  2. Sort them into pause and fallback. Some tasks can wait three hours quite happily. Identify the ones that cannot, usually anything with a fieldwork deadline, a client presentation or a live participant commitment attached.
  3. Keep your data and your working files outside the model. Original data, transcripts, coding frames and editable drafts should live somewhere you control, in formats you can open without the tool that made them.
  4. Preserve a manual route for urgent participant communication. Panel members and interview participants need to hear from a human when systems fail.
  5. Where the risk justifies it, maintain a tested alternative. A second account that nobody has ever used is a comfort rather than a control. Run a real task through the alternative occasionally so you know it works, and so somebody on the team knows how to drive it.

The concise assessment

Capability is rising faster than organisational resilience. That gap is where the risk sits.

Insights teams should keep using these tools and maintain control of three things while doing so: their data, their evidence, and their ability to carry on without any single model or platform. Thursday was a cheap reminder. The next one might arrive on the morning of a client deadline.