True, there's no reason but that doesn't mean it doesn't happen. Many people (I would argue, rightly) equate degraded service to being out of service.
Made-up Scenario: My cluster can handle XXXk users with an SLA of YYms response time. In degraded mode, I'm only handling XXk users with YYYYms response times.
I'm not meeting my SLA for the remaining number of users so I am, in essence, offline.
As to your specific point of "crashing", look at what happened with 37s. Should the server have crashed? No but there was a bug. The reason you add more capacity in the FIRST place is because the existing number of nodes cannot handle the volume. Depending on any number of bugs, issues or configuration your degraded capacity is for all intents and purposes "crashed".
Made-up scenario #2: A single server in your apache configuration can handle 200k concurrent connections reliably with fast response times. Double that load and response times are so long that various devices on the path are timing out the connections as stale. Apache hasn't crashed but it's not really doing anything.
Fast failure is an accepted best practice. Shit, it's baked into Erlang. Kill the process, start a new one and move on. Depending on the nature of the crash, you're doing nothing but churning processes and not actually servicing requests.
The bigger problem is that people don't design for this type of scenario. Static landing pages. Decoupled services instead of monolithic all-in-one containers. Look at github. That's an awesome example of how to degrade service during an outage. Only certain components are "crashed" because everything is fairly decoupled.
Meanwhile there's a guy over here running 4 apps in the same tomcat container that communicate over memory transport with each other or even if he had the common sense to decouple each app, didn't bother to fail fast and was busy spinning up threads trying to communicate with the rest of the services that he can't actually respond to anything externally.