Back
ChatGPT Spent October 5 'Degraded,' Not Down. That's the Outage Your Monitoring Misses.
On October 5, ChatGPT never went fully dark. It just returned errors for hours while OpenAI's own status page called it 'degraded performance.' That is the exact failure a simple up or down check sails right past.

Jagdish Patil
•
"Degraded" Is Doing a Lot of Work in That Sentence
Here is why this one is worth breaking down even though ChatGPT never fully fell over. "Degraded performance" is the status that exists precisely so a company doesn't have to say "down." It covers everything from "0.1% of requests are slow" to "most of your paying users are getting errors and you can still technically load the homepage." Both are degraded. Only one of them matters to the person staring at an error at their desk.
On October 5 the gap between those two readings was the whole story. OpenAI's status page said the app was reachable and most components were operational. The lived experience for a large slice of users was that the product didn't work. Availability varied by plan tier, by model, by which feature you touched. Pages save was broken for one group. Work Mode was broken for another. Conversations errored for a third. No single check would have caught all of it, and a basic "is the site up?" check would have caught none of it, because the site was up the entire time.
The Check That Would Have Told You "All Good"
Most uptime monitoring works like this: it requests a URL, waits for a response, and calls it a pass if the server answers with a 200 and some HTML. Point that kind of check at chatgpt.com on October 5 and it would have come back green every single time. The page loaded. The server was healthy. The homepage rendered. Pass, pass, pass.
Meanwhile the actual job the product exists to do, generating a reply, was failing. This is the oldest trap in monitoring and it has nothing to do with AI. A WordPress site serves its homepage while the checkout throws a database error. An API returns a 200 with an error object in the body. A status page itself stays up and cheerfully reports "all systems operational" while the system it's reporting on is on fire. We wrote about exactly that when Microsoft's status page said "operational" through an Outlook outage that ran 12 hours, and about Spotify breaking in 173 countries while trackers still showed green. October 5 is the same shape with a different logo on it.
The failure isn't that monitoring doesn't work. It's that most monitoring checks the wrong thing. "Did the front door open?" is not the same question as "did the thing behind the door actually do its job?"
What You'd Have Had to Watch to Catch It
To have known about the October 5 degradation from your own tools, a plain homepage ping was never going to be enough. You'd have needed a check that exercised the real path: send a request that mimics what a user does, and verify the response is a correct response, not just any response. A check that reads the status code and the body, so a 200 wrapped around an error still counts as a failure. A check that runs often enough that "errors for ten hours" doesn't mean "you found out on hour nine."
None of that is exotic. It's the difference between monitoring a URL and monitoring an outcome. For a freelancer or a small agency watching client sites, the practical version is: don't just check that the site responds, check the endpoint that actually earns the money. The login. The checkout. The API route the mobile app depends on. Those are the things that go "degraded" while the marketing homepage sits there looking perfectly fine.
Why This Keeps Being Our Beat
We run SIOPS, so we spend a lot of time looking at how things break, and the pattern across every outage we've written up this year is the same. The official signal and the real signal come apart. The status page lags, or softens, or technically tells the truth while completely missing the point. "Degraded performance" on October 5 was true. It was also useless to anyone who needed to know whether to tell their client something was wrong.
You can't fix OpenAI's status page and you can't fix your client's upstream provider. What you can own is whether your own monitoring is asking the right question and whether it can reach you when the answer is bad. A check against the endpoint that matters, running frequently, that treats a wrong response as a failure, is the only thing that would have flagged October 5 from where you sit. And a quiet dashboard notification about it is worth nothing if it lands on a phone that's face down on a nightstand.
The Takeaway
The scary outages aren't the ones that take the whole site down. Those are loud, everyone notices, the provider posts about it, your client already knows. The expensive ones are the quiet partial failures: the site that's up but broken, the status that says "degraded" while a third of users get errors, the checkout that 500s for ten hours while your homepage check keeps flashing green. ChatGPT just spent a day in that exact state, in public, with a status page that never once said "down."
Monitor the outcome, not the front door. And make sure that when the outcome goes bad, the alert is one you physically cannot sleep through.
SIOPS checks the thing that actually matters in under 60 seconds and sounds a Critical Device Alarm that gets through Do Not Disturb and silent mode. A homepage that loads isn't the same as a product that works. Start free at siops.app.



