Yes, sorry! We're investigating, but my current theory is we got overloaded because I relaxed some of our anti-crawler protections a few days ago. (The reason I did that is that the anti-crawler protections also unfortunately hit some legit users, and we don't want to block legit users. However, it seems that I turned the knobs down too far.) In this case, we also had a secondary failure: PagerDuty woke me up at 5:24…
Crazy that Dang literally manages HN in his sleep! We all knew that but I haven't seen any confirmation before this.
Tell HN: HN was down
291–300 of 343 posts
Re: Tell HN: HN was down
#292Earlier quoted context omitted.
Crazy that Dang literally manages HN in his sleep! We all knew that but I haven't seen any confirmation before this.
failing to manage HN in my sleep is more like it
Re: Tell HN: HN was down
#293Re: Tell HN: HN was down
#294Re: Tell HN: HN was down
#295Earlier quoted context omitted.
We all have our moments, and I personally consider HN to be “best effort”, almost like a volunteer project. I’m not certain I’m correct: but thats the optics I have so my expectations are adjusted to that. So don’t beat yourself up please. When I worked for “SaaS unicorn” we typically had multiple levels of escalation, and acknowledging would have done nothing because the alarm would continue firing until fixed. Not…
Yeah we don't exactly pay to be on HN, not much to complain about. I appreciate everyone who works on HN.
Re: Tell HN: HN was down
#296Re: Tell HN: HN was down
#297Yes, sorry! We're investigating, but my current theory is we got overloaded because I relaxed some of our anti-crawler protections a few days ago. (The reason I did that is that the anti-crawler protections also unfortunately hit some legit users, and we don't want to block legit users. However, it seems that I turned the knobs down too far.) In this case, we also had a secondary failure: PagerDuty woke me up at 5:24…
In a situation like this one, good crisis leadership is essential. dang, HN will help you with tips from vast collected experience (please chip in): 1. Blame: The first thing to do is to point the finger. That doesn't mean analysing the technical issue, which can delay this step and limit your options, but figuring out who is politically easiest to blame. Often, that's the new guy, but outside contractors and vendors…
Comprehensiveness: propose extreme, sweeping solutions, such as a lights-out restart of all services, shutting down all incoming requests, and restoring everything to yesterday's backup. This demonstrates that you are ready to address the problem in a maximally comprehensive way. If someone suggests a config change rollback, or a roll-forward patch, ask them why are gambling company time with localized changes, and ask them why are they willing to gamble company time on technical analysis?
Root Cause Analysis Meeting: spend the entire meeting time rehashing the events, pointing fingers and assigning blame. Be sure to mention how the incident could've been over sooner if you just restarted and rolled back every single thing. Be sure to demonstrate out-of-the-box thinking by discussing unrealistic grandiose solutions. When the time is up, run the meeting over by 30 minutes and force all to stay while realistic solution ideas are finally discussed in overtime. This makes it clear to the team that nothing is more important than this incident's RCA--their time surely is not. If someone asks to tap out to pick their kids up after school, remind them that they are making enough money to call them an Uber.
Alerting: be sure to identify anything remotely resembling leading indicators, and add Critical-level wake-you-up alerts with sensitive thresholds for those indicator. Database exceeding 50% CPU? Critical! Filesystem queue length exceeding 5? Critical! Heap usage over 50%? Critical! 100 errors in one minute on a 100000 requests per minute service? Critical! Single log line indicating DNS resolution failure anywhere in the system? Critical! (What if AWS's DNS is down again?) Service requests rate 10% higher than typical peak? Critical! If anyone objects to such critical alerts, ask them why do they want to be responsible for not preventing the next incident?
Re: Tell HN: HN was down
#298Earlier quoted context omitted.
Yeah we don't exactly pay to be on HN, not much to complain about. I appreciate everyone who works on HN.
We pay with content and with the fact that we attract the talent that eventually ends up powering ycombinator investment rounds.
Re: Tell HN: HN was down
#299Re: Tell HN: HN was down
#300Yes, sorry! We're investigating, but my current theory is we got overloaded because I relaxed some of our anti-crawler protections a few days ago. (The reason I did that is that the anti-crawler protections also unfortunately hit some legit users, and we don't want to block legit users. However, it seems that I turned the knobs down too far.) In this case, we also had a secondary failure: PagerDuty woke me up at 5:24…
Crazy that Dang literally manages HN in his sleep! We all knew that but I haven't seen any confirmation before this.