Live data from Hacker News

HN is up again

news.ycombinator.com

311–320 of 390 posts

Re: HN is up again

#311

HN was down because the failover server also failed: https://twitter.com/HNStatus/status/1545409429113229312 Double disk failure is improbable but not impossible. The most impressive thing is that there seems to be no dataloss, almost whatsoever. Whatever the backup system is, it seems rock solid.

According to this comment: https://news.ycombinator.com/item?id=32024485 each server has a pair of mirrored disks, so it seems we're talking about 4 drives failing, not just 2. On the other hand the primary seems to have gone down 6 hours before the backup server did, so the failures weren't quite simultaneous.

> so it seems we're talking about 4 drives failing, not just 2.

Yes—I'm a bit unclear on what happened there, but that does seem to be the case.

Re: HN is up again

#312
post #301

HN was down because the failover server also failed: https://twitter.com/HNStatus/status/1545409429113229312 Double disk failure is improbable but not impossible. The most impressive thing is that there seems to be no dataloss, almost whatsoever. Whatever the backup system is, it seems rock solid.

> Double disk failure is improbable but not impossible. Were they connected on the same power supply? I had 4 different disks fail at the same time before, but they were all in the same PC... (lightning)

[deleted]

Re: HN is up again

#313

HN was down because the failover server also failed: https://twitter.com/HNStatus/status/1545409429113229312 Double disk failure is improbable but not impossible. The most impressive thing is that there seems to be no dataloss, almost whatsoever. Whatever the backup system is, it seems rock solid.

By second disk failure do they mean that the disks on both the primary and fallback servers failed? Or do they mean that two disks (of a RAID1 or similar setup) in the fallback server failed? The latter is understandable, the former would be quite a surprise for such a popular site. That means that the machines have no disk redundancy and the server is going down immediately on disk failure. The fallback server would…

The disks on both the primary and fallback servers definitely failed. Each was in a RAID setup, but those failed too in both cases.

Re: HN is up again

#315
post #149
post #39

Even accounting for this outage, most other SAAS platforms still can't compete with HN's non-existent SLA. Thank you to everyone who keeps this thing running.

> Even accounting for this outage, most other SAAS platforms still can't compete with HN's non-existent SLA. 8 hours of downtime in a given year is 99.9%, so only three nines. The major SaaS platforms all are basically at least as resilient as this, and most have more stringent SLAs.

Most SaaS platforms don't really measure it. They just "do their best".

Plus, it's hard to quantify many cases because there is hard-down and soft-down (partial interruptions).

Re: HN is up again

#316
post #309
post #307

Earlier quoted context omitted.

You are never going to guess how long the HN SSDs were in the servers... never ever... OK... I'll tell you: 4.5years. I am not even kidding.

Let me narrow my guess: They hit 4 years, 206 days and 16 hours . . . or 40,000 hours. And that they were sold by HP or Dell, and manufactured by SanDisk. Do I win a prize? (None of us win prizes on this one).

These were made by SanDisk (SanDisk Optimus Lightning II) and the number of hours is between 39,984 and 40,032... I can't be precise because they are dead and I am going off of when the hardware configurations were entered in to our database (could have been before they were powered on) or when we handed them over to HN, and when the disks failed.

Unbelievable. Thank you for sharing your experience!

Re: HN is up again

#317

Earlier quoted context omitted.

It might not correlate to quality, but if the information found at a website wasn't valued we wouldn't be constantly pulling the site up, just like facebook users do. I'd argue that this site has a good signal/noise ratio by design and specifically to keep you addicted (where "addicted" means using and constantly returning to the site). This site is just designed to attract people who are put off by the kinds of tric…

> I'd argue that this site has a good signal/noise ratio by design Or so we tell ourselves.

I feel the voting process on HN deserves a lot of credit for the quality of front-page content. I wish a knew more about it, other than karma>500 giving the ability to downvote.

Re: HN is up again

#318

Earlier quoted context omitted.

> Double disk failure is improbable but not impossible. It's not even improbable if the disks are the same kind purchased at the same time.

I learned this principle by getting a ticket for a burnt out headlight 1 week after I replaced the other one.

Anyone familiar with car repair will tell you that if one headlight burns out you should just go ahead and replace both, because of this exact phenomenon. I suppose with LEDs we may not have to worry about it anymore

Re: HN is up again

#319
My first thought was “oh sh*t, they finally added this to the list of time-wasting social media blocked sites” :( It was only when I saw it also didn’t work on my phone, that I realized HN itself was actually down

Re: HN is up again

#320
post #301

HN was down because the failover server also failed: https://twitter.com/HNStatus/status/1545409429113229312 Double disk failure is improbable but not impossible. The most impressive thing is that there seems to be no dataloss, almost whatsoever. Whatever the backup system is, it seems rock solid.

> Double disk failure is improbable but not impossible. Were they connected on the same power supply? I had 4 different disks fail at the same time before, but they were all in the same PC... (lightning)

They were in two mirrors, each mirror in a different server. Each server in different racks in the same row. The servers were on different power circuits from different panels.
Post reply on HN