Live data from Hacker News

UK air traffic control meltdown

jameshaydon.github.io

351–360 of 459 posts

Re: UK air traffic control meltdown

#351
post #339
post #274

Earlier quoted context omitted.

That would actually be pretty bad. As mentioned, W3W is propritary, requires an online connection, and has homonyms. On top of that, you need to enter these waypoints into your aircraft's navigation system - sometimes one letter at a time using a rotary dial. These navigation systems will stay in service for decades. Aviation already uses phonetically pronounceable waypoint names. Typically 5 characters long for RNAV…

If only there was a globally unique set of short two-letter names for every country that could be used as prefixes to enforce uniqueness while still allowing every country to manage their own internal waypoint list. If only.

I'm sure they thought about this at some point. Airports already have a country-code prefix. (For example, airports in the Continental US always start with K.)

For whatever reason, by convention navaids never use a country prefix. Even when it would make sense - the code for San Francisco International Airport is "KSFO", but the identifier for the colocated VOR-DME is just "SFO". (Sometimes this does make a big difference, when navaids are located off site - KCCR vs CCR for Concord Airport vs the off-site Concord VOR-DME, for example.)

It's even worse for NDB navaids, which are often just two letters.

Either way, we're stuck with it because it's baked into aircraft avionics and would be incredibly expensive to change at this point.

Re: UK air traffic control meltdown

#352

This is an interesting engineering problem and I'm not sure what the best approach is. Fail safe and stop the world, or keep running and risk danger? I imagine critical systems like trading/aerospace have this worked out to some degree.

The best approach is to simply print the error to the screen, rather than burying it in a “low level log” which only the software vendor has access to.

They had a four hour buffer until the world stopped, but most of that was pissed away because no one knew what the problem was.

Re: UK air traffic control meltdown

#354

> the backup system applied the same logic to the flight plan with the same result Oops. In software, the backup system should use different logic. When I worked at Boeing on the 757 stab trim system, there were two avionics computers attached to the wires to activate the trim. The attachment was through a comparator, that would shut off the authority of both boxes if they didn't agree. The boxes were designed with:…

Wouldn't trim be an number of which a significant tolerance is permissible at any given time? Or does "agree" mean "within a preset tolerance"?

Re: UK air traffic control meltdown

#356
post #335

> the backup system applied the same logic to the flight plan with the same result Oops. In software, the backup system should use different logic. When I worked at Boeing on the 757 stab trim system, there were two avionics computers attached to the wires to activate the trim. The attachment was through a comparator, that would shut off the authority of both boxes if they didn't agree. The boxes were designed with:…

This would have been a 2oo2 system where the pilot becomes the backup. 2oo2 systems are not highly available. Air traffic control systems should at least be 2oo3[1] (3 systems independently developed of which 2 must concur at any given time) so that a failure of one system would still allow the other two to continue operation without impacting availability of the aviation industry. Human backup is not possible becaus…

> Air traffic control systems should at least be 2oo3... Human backup is not possible because of human resourcing and complexity.

But this was a 1oo1 system, and the human backup handled it well enough: a lot of people were inconvenienced, but there were no catastrophes, and (AFAIK) nothing that got close to being one.

As for the benefits of independent development: it might have helped, but the chances of this being so are probably not as much as one would have hoped if one thought programming errors are essentially random defects analogous to, say, weaknesses in a bundle of cables; I had a bit more to say about it here:

https://news.ycombinator.com/item?id=37476624

Re: UK air traffic control meltdown

#357

Earlier quoted context omitted.

First thought that came to my mind as well when I read it. This failover system seems to be more designed to mitigate hardware failures than software bugs.

I also understand that it is impractical to implement the ATC system software twice using different algorithms. The software at least checked for an illogical state and exited, which was the right thing to do. A fix I would consider is to have the inputs more thoroughly checked for correctness before passing them on to the ATC system.

I wonder where most of the complexity lies in ATC. Naively you’d think there would be some mega computer needed to solve the puzzle but the UK only sees 6k flights a day and the scale of the problem, like most things in the physical world, is well bounded. That’s about the same number of buses in London, or a tenth of the number of Uber drivers in NYC.

It would be interesting to actually see the code.

Re: UK air traffic control meltdown

#358
post #344

Earlier quoted context omitted.

not stronger isolation between different flight plans? it seems "obvious" to me that if one flight plan is causing a bug in the handling logic, the system should be able to recover by continuing with the next flight plan and flagging the error to operators to impact that flight only

I'm no aviation expert, but perhaps with waypoints: A B C D E / F G H I J If flight plan #1 is known to be going from F-B at flight level 130, and you have a (supposedly) bogus flight plan #2, they can't quite be sure if it might be going from A-G at flight level 130 at the same time and thus causing a really bad day for both aircraft. I'd worry that dropping plan #2 into a queue for manual intervention, especially i…

The ATC system handled well enough (i.e. no disasters, and AFAIK, no near misses) something much more complicated than one aircraft showing up with no flight plan: the failure of this particular system put all the flights in that category.

I mentioned elsewhere that any ATC system has to be resilient enough to handle things like in-flight equipment failure, medical emergencies, and the diversion of multiple aircraft on account of bad weather or an incident which shuts down a major airport.

As for why the system "pulled the plug", the author of the article suspects that this particular error was regarded as something that would not occur unless something catastrophic had caused it, whereas, in reality, it affected only one flight and could probably have been easily worked around if the system had informed ATC which flight plan was causing the problem.

Re: UK air traffic control meltdown

#359

Earlier quoted context omitted.

I thought that is what the "ICAO pronunciation" was for? "Fly direct Quebec Xray Kilo Charlie Delta"

It is, but fixes are almost always spoken as words rather than letter-by-letter. For this reason, they are usually chosen to be somewhat pronounceable, and occasionally you even get jokes in the names. Likewise, radio beacons and airports are usually referred to by the name of their location; for instance "proceed direct Dover" rather than "proceed direct Delta Victor Romeo". I think a lot of pilots and air traffic c…

I really enjoy the joke names.

Portsmouth, NH has a Sylvester/Tweety Bird approach: ITAWT, ITAWA, PUDYE, TTATT, followed by IDEED for the missed approach.

https://www.pilotsofamerica.com/community/threads/unique-way...

Australia has WALTZ, INGMA, TILDA, and also WONSA, JOLLY, SWAGY, CAMBS, BUIYA, BYLLA, BONGS

https://www.cntraveler.com/stories/2015-06-02/a-pilot-explai...

Disney has a whole lot of special fixes in Orlando and Anaheim. The PIGLT arrival passes through HKUNA, MTATA, JAZMN, JAFAR, RFIKI, TTIGR. I'm fairly sure I've heard about some variants on MICKY, MINEE, GOOFY, PLUTO, etc.

https://aerosavvy.com/wp-content/uploads/2016/04/MCO-PIGLT-S...

According to the same article, Louisville has LUUKE – IAMUR – FADDR.

Re: UK air traffic control meltdown

#360
It must suck to be responsible for a system that everyone depends on and millions of dollars are riding on so you are very reluctant to change it, even if you know it needs technical improvements.

Formal verification or fuzzing could have helped them over that mistrust, but are not panaceas

Post reply on HN