Upstream report · September 2026

625 incidents from 13 status pages: what they say, and what they can't

We mirror the public status pages of the providers our customers depend on and keep every incident they post. This is the first look at what's in the pile.

Key findings, as of 25 September 2026
  • In 588 of 625 incidents the recorded start time is within a minute of the provider's first post. A status page's clock starts when the provider says so.
  • 62 incidents, one in ten, were posted already resolved.
  • Outages keep office hours: 111 posted incidents per weekday against 35 per weekend day. Thursday has 143, Sunday 31. The peak hour is 18:00 UTC.
  • Once admitted, the median incident is resolved in 79 minutes. 134 of 625 were labelled major or critical by the provider.
  • Of 625 incidents, one carries a postmortem.
  • Vercel posted a major incident on 18 September; about 1,700 evaluations of monitors that depend on Vercel, from three continents, saw no failed probe. It was the build pipeline, not serving.

What we did

Repose holds a page when the cause of a failure is upstream. To do that the worker has always polled fourteen public status pages once a minute and kept the current state in memory. Until this week it threw the history away. Now it keeps it: every incident each provider posts, with the full timeline of updates, mirrored every five minutes.

Each provider's feed exposes its 50 most recent incidents, so on day one we had 50 from each. One of the fourteen pages turned out to be another one in disguise: SendGrid's status page is a redirect to Twilio's. So the corpus is 625 incidents from 13 status pages, the oldest from November 2025, all of them in the providers' own words with the providers' own severity labels.

Read every number here with this in mind. Fifty incidents is 16 days of Cloudflare and ten months of Datadog. We don't rank providers by incident count, and neither should you, until each window covers the same quarter. Medians and shares are fair game. Counts aren't.

1. The clock starts when they say it does

In 588 of the 625 incidents, the recorded start time is within a minute of the first public update. That isn't the providers being quick. It's how the status page software works: an incident begins when someone opens it. Whatever happened between the first error and the first post isn't in the record, because the record can't hold it.

Sixty-two incidents, one in ten, were posted already resolved. Start time, end time, and first update, all the same minute. The outage was over before it officially began.

So "median time to resolve" on a status page means time from admission to resolution. The part before admission, the part that woke you up, can only be measured from outside. That's what our probes do, and that's the join we built this week: for every provider incident, what did the monitors that depend on that provider see, in the two hours before the post and until it resolved?

The first result cuts the other way from what we expected. Between 16 and 18 September, Vercel posted four incidents, one of them marked major, "Elevated Errors Triggering Deployments". Our monitors that depend on Vercel were evaluated about 1,700 times inside those windows, from three continents. Not one probe failed. The status page said major. The sites behind it never noticed, because the incident was in the build pipeline, not in serving. A status page describes the provider's day. Your probes describe yours.

2. Outages keep office hours

Sun
31 Mon
87 Tue
94 Wed
123 Thu
143 Fri
108 Sat
39

A weekday averages 111 posted incidents across the corpus. A weekend day averages 35. Thursday alone has more than Saturday and Sunday together, and the peak hour is 18:00 UTC, which is early afternoon on the US East Coast and the end of the day in Europe. The four quietest hours, 02:00 to 05:00 UTC, produce fewer incidents combined than that one hour.

Traffic doesn't look like that. Deploys do. Most of what a status page records is a change someone made, found out about, and rolled back. The lesson for anyone on call is unglamorous: the dangerous window is Thursday afternoon, not Saturday night.

3. Seventy-nine minutes, once admitted

79 min
median time from first post to resolved, all 625
134
labelled major or critical by the provider, 21%
62
posted already resolved

Half of all incidents are closed within 79 minutes of being opened. Who closes fastest, over the incidents we hold:

ProviderIncidents heldSinceMedian fixFixed < 1 h
Zoom5020 Jul 202636 min63%
Netlify5011 Mar 202637 min60%
Stripe506 Mar 202651 min52%
OpenAI2510 Sep 202651 min56%
Vercel5018 May 202655 min56%
Anthropic5022 Jul 202663 min46%
GitHub5024 Jul 202671 min42%
Cloudflare509 Sep 202671 min43%
Datadog505 Nov 202579 min38%
Supabase5030 Jun 20262 h 2 min29%
MongoDB507 Feb 20262 h 7 min31%
DigitalOcean5022 Apr 20262 h 38 min20%
Twilio5016 Sep 20265 h 44 min14%

Twilio's number needs its context: most of its incidents are SMS delivery delays to a single carrier in a single country, which Twilio can post about but can't fix. It's upstream of upstream. DigitalOcean's is more straightforward: infrastructure incidents take longer than software ones.

The long tail is where the labels get strange. MongoDB opened a "minor" incident on 1 March for impaired cluster operations in two Middle East regions and closed it 186 days later. Supabase kept "network access issues affecting a limited number of users in Myanmar" open for 27 days. Both are honest, in a way. Both would also be invisible in any average.

One more thing about labels: Cloudflare has posted 50 incidents since 9 September and called none of them major. Anthropic called 20 of its 50 major. That tells you about two companies' vocabularies, not about two companies' reliability, which is why we show the label and don't rank on it.

4. What actually breaks

Status pages tag incidents with components, and the tags are more useful than the headlines.

5. One postmortem

Of 625 incidents, exactly one carries the status "postmortem": DigitalOcean's 24 August outage of its control panel and API, a critical, 25 hours long, with a write-up attached. Everything else ends at "resolved". The average incident gets 3.9 updates, the median gets 3: investigating, identified, resolved. What went wrong, in the sense of why, is almost never written down where customers can read it.

We're not going to pretend that's easy to fix. We will point out that the one company that did it is not the one with the fewest incidents.

What we do with this

The live version of the table above is at /upstream, computed from the same rows on every request, with each provider's last five incidents linked to the provider's own post. When a monitor of yours declares a dependency on one of these providers and that provider has an open incident, the engine holds the page and writes down why. Every held decision is in your audit trail, next to what the provider said at the time.

This report comes out monthly, from the same query, and the numbers will get better in two ways: the windows will line up as the corpus grows past its first quarter, and the admission gap, the minutes between our probes failing and the provider posting, will start to fill in as more dependent monitors sit behind these providers during real outages. When it does, it'll be here.

What we won't do: rank providers by incident count on mismatched windows, reinterpret their severity labels, or scrape anything they haven't published as data. Fourteen status pages, all on the same platform, is also not the whole cloud. AWS and Azure publish RSS instead, and they're next.

If your site depends on any of these thirteen, Repose can watch it from three continents and stay quiet when the cause is upstream.

Start free How the hold decision works →