Live data from Hacker News

Cause of YC/HN outage discovered

archub.org

101–110 of 110 posts

Re: Cause of YC/HN outage discovered

#103
post #28

Earlier quoted context omitted.

150K hits in 30 minutes is 83 hits/second. That's a lot to ask for from a shared-hosting account. When the traffic starts to impair neighboring sites, something has to be done. Just about any ISP will do the same thing: block the site with the surge, that could possibly make other arrangements, rather than inconvenience other customers whose traffic is as expected/usual. The detail missing so far is why Pair noticed…

So let's see... you're hosting a website that has gotten large; that is, it's grown to the point that it will need higher-cost services in order to meet demand. You have a chance to add a valuable customer to your client base. How best to handle this? a) anything b) except c) killing their service

Yeah, Pair screwed themselves pretty badly by pulling this trick. They would have been far better off doing something like:

Dear [account_contact_user]

Your website traffic has risen beyond the maximum threshold of [threshold_amt] for the [name_of_level] level of service.

Since we appreciate your business of the last [length_of_service], we have given you a 24 hour courtesy upgrade to our next level of service -- [name_of_next_level]. If, by [end_time] you decide to keep this level of service you must contact our sales center to arrange payment. Otherwise we will have to start throttling traffic to your server so that it remains below the threshold of [threshold_amt] and does not impact our other customers.

If you have any questions about this courtesy upgrade, or wish to keep this new level of service, please contact [account_manager] at [account_manager_details].

Thank you for using Pair Networks for your hosting needs.

Re: Cause of YC/HN outage discovered

#104

Earlier quoted context omitted.

How best to handle this? Of course, the flip side is that leaving it running adopts an attitude of "screw all our other customers, they can eat crappy service while we kiss up to the popular guys who are chewing up everybody else's server resources". Which isn't what I'd look for in a hosting provider...

False dichotomy. The correct way to handle this would have been to temporarily move the shared server to hardware where it won't impact other customers and notify YC that they need to move their server to a bigger server or the site will have to be shut down. Presumably with an ultimatum of a week or whatever.

Agreed - the whole reason you pay a host is that there's some level of management and responsibility there. You're not just renting hardware.

Re: Cause of YC/HN outage discovered

#105
post #10

Shared hosting is generally sucky for anything remotely successful. I'm amazed you got away with storing your static assets there so long! When I got 100k hits in a day on my first blog, Dreamhost promptly shut it down without warning (in the middle of a slashdotting!)

Dreamhost? Their main advertising point is UNLIMITED TRAFFIC!!! I guess I'll scratch them off the list of potential hosts.

[deleted]

Re: Cause of YC/HN outage discovered

#106

Earlier quoted context omitted.

How best to handle this? Of course, the flip side is that leaving it running adopts an attitude of "screw all our other customers, they can eat crappy service while we kiss up to the popular guys who are chewing up everybody else's server resources". Which isn't what I'd look for in a hosting provider...

False dichotomy. The correct way to handle this would have been to temporarily move the shared server to hardware where it won't impact other customers and notify YC that they need to move their server to a bigger server or the site will have to be shut down. Presumably with an ultimatum of a week or whatever.

This is exactly the kind of approach that a smart company which cares about business would take.

Re: Cause of YC/HN outage discovered

#107

Kinda funny and sad that so many people here only see this as some sort of stupid or unfair action against HN, seemingly without even acknowledging that every single other customer on that shared server had as much right individually, and more right collectively, to not have their performance negatively impacted by HN. Yeah, it sucks that one of our favorite tech news sites was impacted by this, but how impacted were…

Strawman. Nobody's saying HN/YC should be allowed to overuse resources (then again, I don't know the terms they had agreed to). But not giving a warning is crappy customer service, no matter how many other folks do likewise, how many years Pair has been doing other stuff well, or however else you want to spin it. In this case, it also happens to be a big sales screwup.

Well, you didn't identify what I said that you think is a strawman, but I'll point out yours.

Fine, nobody's saying HN/YC should be allowed to overuse resources. I didn't say anybody was saying that.

Not giving a warning probably doesn't count as great customer service, but then again, once the problem had been identified by Pair, and once they knew of the negative impact HN was having on every other paying customer on that server, what kind of customer service to those other customers would it have been for Pair to fire off an email to HN then wait an hour, or thirty minutes, or ten minutes, before shutting it down?

How long should Pair have allowed HN to impact other customers to satisfy folks here? And what makes HN more important than any other paying customer on that server?

Oh right, it's because you read and like HN, which, ironically, so do I.

As for it being a sales screwup, maybe. I kinda doubt there is a great deal of overlap between HN readership and the average potential Pair customer. We could also suggest that Pair taking action to protect all those other customers on the server is an example of how they would provide good service to the many when they're being hammered by one overpowering fellow customer.

Re: Cause of YC/HN outage discovered

#109
post #96

I typically don't tell other people how to run their businesses, but if a similar issue brought my website down and I were to post about the cause s , I might focus more on my failures in capacity planning, vendor selection, and monitoring rather than on my vendor's lackluster customer service. User-visible failures are, ultimately, process failures on my part, regardless of the surface cause. A nice side effect of t…

It didn't bring HN down. HN deliberately didn't rely on that server for anything except hosting static content that was also duplicated on this server. I planned in advance for the possibility that the other server wouldn't be usable, by writing the code so that I could switch to serving the same content off news by changing one variable, which I did. As a result service was barely affected.

In short, Pair flaked, but we had in fact planned the system in a way that protected us against it.

Re: Cause of YC/HN outage discovered

#110
post #28

Earlier quoted context omitted.

150K hits in 30 minutes is 83 hits/second. That's a lot to ask for from a shared-hosting account. When the traffic starts to impair neighboring sites, something has to be done. Just about any ISP will do the same thing: block the site with the surge, that could possibly make other arrangements, rather than inconvenience other customers whose traffic is as expected/usual. The detail missing so far is why Pair noticed…

So let's see... you're hosting a website that has gotten large; that is, it's grown to the point that it will need higher-cost services in order to meet demand. You have a chance to add a valuable customer to your client base. How best to handle this? a) anything b) except c) killing their service

What about the dozen other customers who are calling you to complain about crappy service right now?
Post reply on HN