Live data from Hacker News

NYSE Tuesday opening mayhem traced to a staffer who left a backup system running

bloomberg.com

231–233 of 233 posts

Re: NYSE Tuesday opening mayhem traced to a staffer who left a backup system running

#231
post #229

Earlier quoted context omitted.

I used to work for a small startup, and postmortems were truly no blame - engineers would talk about exactly what happened and wouldn't hesitate to put the blame on their mistakes. But as the company grew, the postmortems became more about blame since now you're not blaming an engineer, but an entire team so singling them out isn't personal. The postmortems were no longer a single engineer describing what happened in…

This happens within big organizations that are large enough where they start having that internal small company feel within units. I would say a good program, which could be small or a chunk in a massive org, does a blameless post mortem. A few years back, a task to modify an index was given to a scrum team. The lead was away and the senior people could not be bothered. The junior developer stack overflowed an answer…

Suddenly I don't feel so bad about deleting an entire PVCS repository (happily answering 'yes' to all the 'are you sure?' questions) at 4:30PM on a Friday.

Re: NYSE Tuesday opening mayhem traced to a staffer who left a backup system running

#232
post #211

Earlier quoted context omitted.

I had a project that was really important once. I made a tiny mistake that had a big consequence - a couple of hours of potential lost revenues from our customers. I fixed my mistake with both my boss and CEO nearby. I said after I pushed the fix that I really need more resources around it. That little light of "yeah, this is important" that should have flickered didn't. :) I will not be surprised if nothing gets fix…

But how much actual lost customer revenue? Also, did the customer even notice or not? You're reminding me of the difference between engineers and non-technical managers; to many of the latter something's only a problem if/when the customer or senior mgmt are on the phone complaining about it. Until then it's all naysaying engineers being too pessimistic about process and risk.

No one gave me a figure, but the product going down does have a direct impact on customer revenues. So, while no one actually gave me direct numbers, it was felt quite a bit. The most notable thing from the incident was where certain account managers in Japan had to do formal mea culpas because of this mistake. So, in other words, they were on calls, but I got back just "you can't have this happen again - do something." Could someone else at least do code review? I was depressed that I got a lot of "well, we don't know the code base" - the issue was lost in the shuffle also with management. So, I just grew more pessimistic about process and risk. It's not healthy and I would not bottle it up now, but I did, and that was my mistake. I ended up hating that project, but at least I got to train someone else to work on it (who has a team around him.)

Re: NYSE Tuesday opening mayhem traced to a staffer who left a backup system running

#233

Earlier quoted context omitted.

Yeah my point is only that it's a full DR site, not a "backup" that was left running as the article pointed out (and as a lot of commenters are insinuating).

Ah, I see. Some quick internet searching shows that from time to time (rarely), NYSE operates the Cermak site as the primary site for at least part of their operations - for up to a week at a time. It seems to have multiple purposes - DR, customer software testing, etc.

Well, you do that to make sure it works. At one of my previous jobs, we'd switch the primary and secondary datacenters every six months just to make sure we could.
Post reply on HN