- In 588 of 625 incidents the recorded start time is within a minute of the provider's first post. A status page's clock starts when the provider says so.
- 62 incidents, one in ten, were posted already resolved.
- Outages keep office hours: 111 posted incidents per weekday against 35 per weekend day. Thursday has 143, Sunday 31. The peak hour is 18:00 UTC.
- Once admitted, the median incident is resolved in 79 minutes. 134 of 625 were labelled major or critical by the provider.
- Of 625 incidents, one carries a postmortem.
- Vercel posted a major incident on 18 September; about 1,700 evaluations of monitors that depend on Vercel, from three continents, saw no failed probe. It was the build pipeline, not serving.
What we did
Repose holds a page when the cause of a failure is upstream. To do that the worker has always polled fourteen public status pages once a minute and kept the current state in memory. Until this week it threw the history away. Now it keeps it: every incident each provider posts, with the full timeline of updates, mirrored every five minutes.
Each provider's feed exposes its 50 most recent incidents, so on day one we had 50 from each. One of the fourteen pages turned out to be another one in disguise: SendGrid's status page is a redirect to Twilio's. So the corpus is 625 incidents from 13 status pages, the oldest from November 2025, all of them in the providers' own words with the providers' own severity labels.
1. The clock starts when they say it does
In 588 of the 625 incidents, the recorded start time is within a minute of the first public update. That isn't the providers being quick. It's how the status page software works: an incident begins when someone opens it. Whatever happened between the first error and the first post isn't in the record, because the record can't hold it.
Sixty-two incidents, one in ten, were posted already resolved. Start time, end time, and first update, all the same minute. The outage was over before it officially began.
So "median time to resolve" on a status page means time from admission to resolution. The part before admission, the part that woke you up, can only be measured from outside. That's what our probes do, and that's the join we built this week: for every provider incident, what did the monitors that depend on that provider see, in the two hours before the post and until it resolved?
The first result cuts the other way from what we expected. Between 16 and 18 September, Vercel posted four incidents, one of them marked major, "Elevated Errors Triggering Deployments". Our monitors that depend on Vercel were evaluated about 1,700 times inside those windows, from three continents. Not one probe failed. The status page said major. The sites behind it never noticed, because the incident was in the build pipeline, not in serving. A status page describes the provider's day. Your probes describe yours.
2. Outages keep office hours
A weekday averages 111 posted incidents across the corpus. A weekend day averages 35. Thursday alone has more than Saturday and Sunday together, and the peak hour is 18:00 UTC, which is early afternoon on the US East Coast and the end of the day in Europe. The four quietest hours, 02:00 to 05:00 UTC, produce fewer incidents combined than that one hour.
Traffic doesn't look like that. Deploys do. Most of what a status page records is a change someone made, found out about, and rolled back. The lesson for anyone on call is unglamorous: the dangerous window is Thursday afternoon, not Saturday night.
3. Seventy-nine minutes, once admitted
Half of all incidents are closed within 79 minutes of being opened. Who closes fastest, over the incidents we hold:
| Provider | Incidents held | Since | Median fix | Fixed < 1 h |
|---|---|---|---|---|
| Zoom | 50 | 20 Jul 2026 | 36 min | 63% |
| Netlify | 50 | 11 Mar 2026 | 37 min | 60% |
| Stripe | 50 | 6 Mar 2026 | 51 min | 52% |
| OpenAI | 25 | 10 Sep 2026 | 51 min | 56% |
| Vercel | 50 | 18 May 2026 | 55 min | 56% |
| Anthropic | 50 | 22 Jul 2026 | 63 min | 46% |
| GitHub | 50 | 24 Jul 2026 | 71 min | 42% |
| Cloudflare | 50 | 9 Sep 2026 | 71 min | 43% |
| Datadog | 50 | 5 Nov 2025 | 79 min | 38% |
| Supabase | 50 | 30 Jun 2026 | 2 h 2 min | 29% |
| MongoDB | 50 | 7 Feb 2026 | 2 h 7 min | 31% |
| DigitalOcean | 50 | 22 Apr 2026 | 2 h 38 min | 20% |
| Twilio | 50 | 16 Sep 2026 | 5 h 44 min | 14% |
Twilio's number needs its context: most of its incidents are SMS delivery delays to a single carrier in a single country, which Twilio can post about but can't fix. It's upstream of upstream. DigitalOcean's is more straightforward: infrastructure incidents take longer than software ones.
The long tail is where the labels get strange. MongoDB opened a "minor" incident on 1 March for impaired cluster operations in two Middle East regions and closed it 186 days later. Supabase kept "network access issues affecting a limited number of users in Myanmar" open for 27 days. Both are honest, in a way. Both would also be invisible in any average.
One more thing about labels: Cloudflare has posted 50 incidents since 9 September and called none of them major. Anthropic called 20 of its 50 major. That tells you about two companies' vocabularies, not about two companies' reliability, which is why we show the label and don't rank on it.
4. What actually breaks
Status pages tag incidents with components, and the tags are more useful than the headlines.
- Stripe: 32 of 50 incidents touched "Acquirers and payment methods". Elevated Klarna errors, elevated Visa declines, one card network at a time. Almost none touched the API itself. If you monitor stripe.com you'll see nothing. Monitor your own checkout.
- Vercel: 16 of 50 touched "Builds". The pipeline, not the edge. Which is the 1,700-evaluations story from above, told from the other side.
- Datadog: the component that appears most often, 23 of 50 incidents, is "Monitors". The second, at 18, is "Metrics and Infra Monitoring". A monitoring vendor's status page is mostly about monitoring being delayed, which is the one failure a monitoring vendor's own alerts can't tell you about.
- Anthropic: 38 of 50 touched claude.ai, 33 the API. When it goes, it mostly goes together.
5. One postmortem
Of 625 incidents, exactly one carries the status "postmortem": DigitalOcean's 24 August outage of its control panel and API, a critical, 25 hours long, with a write-up attached. Everything else ends at "resolved". The average incident gets 3.9 updates, the median gets 3: investigating, identified, resolved. What went wrong, in the sense of why, is almost never written down where customers can read it.
We're not going to pretend that's easy to fix. We will point out that the one company that did it is not the one with the fewest incidents.
What we do with this
The live version of the table above is at /upstream, computed from the same rows on every request, with each provider's last five incidents linked to the provider's own post. When a monitor of yours declares a dependency on one of these providers and that provider has an open incident, the engine holds the page and writes down why. Every held decision is in your audit trail, next to what the provider said at the time.
This report comes out monthly, from the same query, and the numbers will get better in two ways: the windows will line up as the corpus grows past its first quarter, and the admission gap, the minutes between our probes failing and the provider posting, will start to fill in as more dependent monitors sit behind these providers during real outages. When it does, it'll be here.
What we won't do: rank providers by incident count on mismatched windows, reinterpret their severity labels, or scrape anything they haven't published as data. Fourteen status pages, all on the same platform, is also not the whole cloud. AWS and Azure publish RSS instead, and they're next.
If your site depends on any of these thirteen, Repose can watch it from three continents and stay quiet when the cause is upstream.
Start free How the hold decision works →