Live data from Hacker News

Level 3 technician's misstep causes largest telephone outage ever reported

fiercetelecom.com

81–85 of 85 posts

Re: Level 3 technician's misstep causes largest telephone outage ever reported

#81
post #12

Earlier quoted context omitted.

An "Are you sure?" prompt goes a long way and sometimes takes a really long time to get added, even when it's historically been a problem[0]. Particularly in these kinds of cases -- this is software that is used by very few people at very few companies and at those companies, it's used very rarely. That nobody at Level 3 knew leaving that field blank would cause that issue doesn't surprise me at all. We had managemen…

While "Are you sure?" is often a great prompt, there are two things that can be done to improve it. 1) Just putting AYS? after every prompt is bad, people stop taking the time to think and just mash Y. 2) Discussing the consequences in the AYS? is better. For instance if you @channel in slack you get an AYS that lists the number of people you are going to piss off with the notification. Great UX... and yet people at…

I think a lot of those problems can be solved with a few techniques:

In the case of the famous rm -rf /, having a flag to ditch confirmation (which, IIRC, the reason the confirmation prompt was not there in the first place is because "f" means force, but hey). I think about "zypper" (openSUSE's package manager), which has the -y flag for installs/updates, one to auto-accept licenses and one to force resolving conflicts aggressively so that there are, literally, no prompts. To me, having a -y and an auto-accept licenses is unnecessary, but I suspect that may exist for legal reasons. Having the extra flag for aggressively accepting conflicts is a really good idea because in interactive mode you're often given 3 choices and often all three of those choices will break something. Conflicts/package resolution issues aren't common and when the creep up it's usually because you've got a unique configuration and need to do some other steps before you're going to get a successful installation.

As for the "Are you sure?" GUI prompts where flags won't cut it, providing a "[ ] Never ask again" is usually suitable. Though in the case of Slack, I'm with you. My answer to that would be after a person has used @channel more than a twice in a day to throw another prompt up that says "You're going to get a reputation for being an obnoxious dick. Knock it off." ...and that's why I'm not a UX designer.

Re: Level 3 technician's misstep causes largest telephone outage ever reported

#82
post #30
post #10

Earlier quoted context omitted.

I commented more extensively in the root of the post, but you can't even begin to imagine. Think about every script you've ever written for "some thing at home" and how you only cared that it worked for the very narrow, specific, circumstances you were looking for. Maybe you left out error handling and just let it crash when you failed to put in the right parameter. Who cares? It's just a script for your one, lonely,…

> Think about every script you've ever written for "some thing at home" and how you only cared that it worked for the very narrow, specific, circumstances you were looking for. Maybe you left out error handling and just let it crash when you failed to put in the right parameter. Who cares? It's just a script for your one, lonely, workstation/server. I find seeing this mentioned oddly comforting. I write my worst soft…

I joke that I have an ever growing private repository of code I'm too embarrassed to publish publicly. Only partly joking; it's actually several repositories. But as others have said, people who are programmers more than just professionally do this all the time. There's very few things that involve using a computer that I run into day-to-day that I don't think, "I could do this faster with a dirty script[0]"

I used to do interviews and would often encounter candidates who had no public repositories, anywhere, GitHub or otherwise. I learned quickly to be very disarming before even asking and settled on something along the lines of "Look, I know when you write things for yourself, you're doing it with very limited time and for an audience who is more interested in it doing 'The Thing' without regard for anything resembling best practices, or even typical practices. I fully expect lousy code and that's perfectly fine, but I really need as many recent samples as you can possibly give me before this evening to be prepared.[1]"

What wasn't probably realized by the candidates was that if I'd made that phone call they were getting the technical interview regardless of the code quality and I was barely going to glance at their code until the interview. And the best thing they could do for me was to give me code that was on the bad side of things[2]. The first question I'd ask is "OK, imagine you have all of the time you'd ever need to make this perfect. The goal is to get there in stages and maximize the improvements as early as possible. What would you change and in what priority." There's no way in their development career that they're not going to encounter something as awful as they've written in production and be asked to fix it, and they're going to have to start with "make it work again" and in a limited number of stages "get it to the best state; hopefully to an ideal state" but the goal is to get to as stable of a state as possible before you're pulled somewhere else.

[0] Or better, stop doing it at all by putting that dirty script in a cron job that won't have any logging, I'll probably be really happy with for the first few weeks, then forget it was there and not realize a few months later when it stops working.

[1] It was kind of a dick move and I always felt a little guilty, but I know that my instinct would be to spend an evening cherry picking the best examples which I'd then go without sleep to clean up as best I could before morning. And it wasn't a situation where I wanted to judge them at their worst "Well, if their worst code looks this good, they must be good."; if I got a code sample that was too good, I figured it was the rare script that was cared about but that the developer didn't want to unleash on the world and feel like they had to support. When this happened, it was always from a candidate who gave me one maybe two things while claiming to be a code addict and I always asked those candidates to rectify their love for software with only having a couple of very small, albeit well-written code samples. I recommended one guy who admitted he had a lot of other code but was too embarrassed by it and then logged into his BitBucket repo.

[2] Which I found worked best when I simply just asked candidates for their worst code and told them, generally, that I'd like to get a feel for how they handle eliminating technical debt. If the code was too good, I'd have to find something that I felt they should be able to quickly understand well enough to offer ideas for improvements; that rarely worked out well -- it's always easier on the candidate when it's code they've written because they're, at lease possibly, likely to be familiar with it.

Re: Level 3 technician's misstep causes largest telephone outage ever reported

#83
post #74
post #18

Earlier quoted context omitted.

Three things. 1. Thanks for the awesome TIL 2. Rumors fly about old versions of Windows, OS/2, etc still being actively used. I like to pin down and file away usage/year correlations, where possible. What sort of timeframe (roughly) was Win98 in active use here? 3. Regarding [1], I have an ancient [runs downstars to check] Compaq Prosignia 300 server here and I discovered in the (DR-DOS based) BIOS at one point that…

1. You're kindly welcome. 2. The last I had heard about that machine, specifically, was around 2011 and I'm fairly certain it was there when they eliminated the NOC in Detroit which was around 2013, if memory serves. The thing is, I would be surprised if it was actually gone. 3. Nice - I was well known was the guy who could fix anything over at Level 3 around the IT side of the house. Around 2014, I was asked to take…

Interesting... I think you just helped me figure something (admittedly rather simple) out. Quite trivial (not nearly as interesting as your experience), but mildly related.

Many years ago I happened to find an ancient-looking machine buried in a spare room at a church. I think the room was occasionally used as an ad-hoc creche area.

After finally locating an IEC cable for it and finally getting it to boot, I found that it was of the opinion it didn't have any HDDs attached.

So, I went into the BIOS, and - yes! Just old enough to require manual CHS configuration, but juuust new enough to have a manual autodetect routine!

Turned out it was a cute 25MHz-or-so (IIRC) 486 with something like a 200MB HDD. Had some demos on it that I've long forgotten the names of. Was fun to find that machine.

Despite being so trivial that no conveniently-placed homemade rescue labels were needed, the people there also wondered how I'd figured out what was wrong with it as well.

(Honestly, I really want to work somewhere everybody's still using ancient equipment. Partly because it's what I've been exposed to for most of my life and I really like it, and partly because I'm still yet to have the chance to acclimatize to newer stuff and almost all of the small bit of knowledge I've accumulated covers older tech.)

Regarding what you helped me figure out, I've been wondering for years why that BIOS decided it didn't have a HDD. Initially I thought the HDD was on the way out and decided not to show up one day, and the BIOS happily deregistered it [after someone saw an indecipherable POST error and hit F1 or whatever]. Now I wonder if maybe the rechargeable battery went flat after the machine was left off for ages, then someone turned it on, hit F1 to accept the "bad CRC" error (I never saw one) and didn't know to do an autodetect for the HDD. It's possible. I know some BIOSes remember "time wasn't re-set after fail"; I never saw a time POST error, and don't remember if I checked the clock to see if it was wrong.

Re: Level 3 technician's misstep causes largest telephone outage ever reported

#84
post #80
post #78

Earlier quoted context omitted.

Nice. I'm guessing I can't read the source to this particular fault tree, but I wonder where I might find others. Preferably without digging through e.g. troves of court documents and the like.

https://www.ntsb.gov/investigations/AccidentReports/Pages/Ac...

Ooh, interesting. Thanks!

And I'd somehow gotten in my head the NTSB were only aviation, probably from old TV shows. TIL about their actual name.

Re: Level 3 technician's misstep causes largest telephone outage ever reported

#85
post #44

Well, I dunno. When I worked at the Telco that serves all of Northern Canada - the telco that has the largest operating area of any telco in the world, in fact - we had an outage that took out everything for between 1 and 3 days. When I say everything, I mean if you picked up your phone you didn't get a dial tone. Or cell phone. Or internet. Even people that still have 2-way radios to use as phones were out of luck.…

By population the entire area served by northwestel is smaller than a single rural wa state county, however.
Post reply on HN