Due to failure in the IT system, it is not possible to run any trains today
1–10 of 365 posts
Re: Due to failure in the IT system, it is not possible to run any trains today
#2> The IT failure occurred at the end of the morning. It affected the system that generates up-to-date schedules for trains and staff. This system is important for safe and scheduled operations: if there is an incident somewhere, the system adjusts itself accordingly. This was not possible due to the failure.
Re: Due to failure in the IT system, it is not possible to run any trains today
#3I do hope they have a more detailed RCA at some point, right now this is the only helpful paragraph in the article about what happened: > The IT failure occurred at the end of the morning. It affected the system that generates up-to-date schedules for trains and staff. This system is important for safe and scheduled operations: if there is an incident somewhere, the system adjusts itself accordingly. This was not pos…
If I had to guess, it probably was an issue with scheduling/timetable software rather than anything to do with the trains, rails, etc. Nothing exceedingly seriously or difficult to correct.
Re: Due to failure in the IT system, it is not possible to run any trains today
#4Re: Due to failure in the IT system, it is not possible to run any trains today
#5tl;dr : Restarting train schedule mid-day is too hard and unsafe. So we're sending them back to depot instead.
Re: Due to failure in the IT system, it is not possible to run any trains today
#6tl;dr : Restarting train schedule mid-day is too hard and unsafe. So we're sending them back to depot instead.
Why "unsafe"? Isn't safety supposed to be handled at a different abstraction level (in hardware)?
Other systems that might normally provide safety critical redundancy could be providing the sole measure of safety, with no other redundancy available in case one of those fails.
“Unsafe” is always defined based on context.
Re: Due to failure in the IT system, it is not possible to run any trains today
#7tl;dr : Restarting train schedule mid-day is too hard and unsafe. So we're sending them back to depot instead.
Why "unsafe"? Isn't safety supposed to be handled at a different abstraction level (in hardware)?
Re: Due to failure in the IT system, it is not possible to run any trains today
#8Also probably worth noting is that trains here are used fairly frequently, however unfortunately this isn't the first time trains are stopping - for example snow is a common reason for delayed/reduced service. (Snow isn't very common here in the NLs)
Re: Due to failure in the IT system, it is not possible to run any trains today
#9Earlier quoted context omitted.
Why "unsafe"? Isn't safety supposed to be handled at a different abstraction level (in hardware)?
Defense in depth? With this safety feature inoperable, the margin of safety is reduced. Other systems that might normally provide safety critical redundancy could be providing the sole measure of safety, with no other redundancy available in case one of those fails. “Unsafe” is always defined based on context.
Edit: and to be clear, we're not talking about holding down a resetable circuit breaker to avoid being late, we're talking about the rail network of an entire nation being inoperable for a day.
Re: Due to failure in the IT system, it is not possible to run any trains today
#10Earlier quoted context omitted.
Defense in depth? With this safety feature inoperable, the margin of safety is reduced. Other systems that might normally provide safety critical redundancy could be providing the sole measure of safety, with no other redundancy available in case one of those fails. “Unsafe” is always defined based on context.
What's the point of having redundancy if you can't use it to avoid an impact to service? Edit: and to be clear, we're not talking about holding down a resetable circuit breaker to avoid being late, we're talking about the rail network of an entire nation being inoperable for a day.
You can never know if the primary safety system is functioning perfectly, so you need other systems to be there to step in when the primary fails unexpectedly.
If you detect the primary system has failed, isn’t it reasonable that you should stop operation as quickly and safely as possible, and be thankful nothing bad happened while you lacked redundancy? Any SPoF could be fatal for hundreds of people.