Live data from Hacker News

Fog Creek is about to go down

fogcreekstatus.typepad.com

71–80 of 193 posts

Re: Fog Creek is about to go down

#71
post #24
post #18

Earlier quoted context omitted.

Wow you expect fog creek to be down for that long when it does? Why can't you just scp everything on there to your other DC? or at least just move the harddrives? Seems like it would cost less time.

Kiln alone has >4 TB of data; you want to SCP that with 90 minutes heads-up? Having power and/or rack space is not the same as having servers, switches, etc. anyway. Hopefully it will not be down that long. We'll let you know more when we know more.

You could have rsync'd it with a few days heads up, and freshened that in the last 90min.

Re: Fog Creek is about to go down

#72
post #13

"Given the preparation work that's gone into this, we are confident that all of our services will remain available to our customers throughout the weather." - yesterdays update. Try not to let your fingers type cheques your datacenters can't cash...!

Nobody was really expecting this much of Manhattan to lose power.

It's pretty amusing to see you posting this when a couple days ago, you were accusing the media of "blowing it out of proportion". Oh hey, it turns out they were right.

http://news.ycombinator.com/item?id=4706959

Re: Fog Creek is about to go down

#73
post #7

Too bad that the backup generator refuelling pumps have been submerged (while the generators themselves are running). That sounds like some really ... unfortunate planning of the positioning of these machines, made me think of the Daiichi incident when backup assets failed to come online because parts of the backup infrastructure were destroyed. Not as serious, of course. Fog Creek's hosting isn't a nuclear power pla…

New York has some of the strictest building codes going, and I wouldn't be surprised if that played a part.

Note that the pumps were in the basement because that's where the refueling location was, which logically it would be for a variety of practical reasons (you can't drive a fuel truck up an escalator, etc.)

Re: Fog Creek is about to go down

#74
post #24
post #18

Earlier quoted context omitted.

Wow you expect fog creek to be down for that long when it does? Why can't you just scp everything on there to your other DC? or at least just move the harddrives? Seems like it would cost less time.

Kiln alone has >4 TB of data; you want to SCP that with 90 minutes heads-up? Having power and/or rack space is not the same as having servers, switches, etc. anyway. Hopefully it will not be down that long. We'll let you know more when we know more.

Do you guys do off-site backups?

Re: Fog Creek is about to go down

#75
post #54
post #7

Too bad that the backup generator refuelling pumps have been submerged (while the generators themselves are running). That sounds like some really ... unfortunate planning of the positioning of these machines, made me think of the Daiichi incident when backup assets failed to come online because parts of the backup infrastructure were destroyed. Not as serious, of course. Fog Creek's hosting isn't a nuclear power pla…

That's the first thing I thought of too. Seems like there should be a new rule that your backup generator infrastructure should be on at least the second floor.

That means that the fuel tanks will have to be there as well or you're going to need some specialized lines/equipment to prime the pumps. How much would a 3" cylinder of diesel fuel about ten foot tall weigh? (What ever it is that would be a crap load of vacuum to produce/maintain). Then you have the weight of the tanks themselves and refilling logistics. All fun engineering problems. :)

Re: Fog Creek is about to go down

#76
post #45
post #32

Earlier quoted context omitted.

Why would moving the servers be hard? You can fit 100 terabytes of storage in a shoe-box these days. I'd be extremely surprised if you couldn't run all of FogCreek off of a single 10U blade enclosure. That would be up to 128 CPU cores; I suspect they need only a small fraction of that. On the other hand, with a 1000/Mbit uplink that they were allowed to saturate, they'd still only be able to copy out 1 terabyte in 3…

Depends on your mental model of "servers." People have different views of servers from "that nosiy hot thing under my desk" to multiple cages and dozens of racks with 1G, 10G, 40G fiber, rack switches, core switches, edge switches, terminal servers, database servers, application servers, monitoring servers, corporate boxes, .... All with various weights, accessibility, cable routing, and those damn three servers with…

In this case my mental model of "servers" is "the computers that run the specific small company under discussion, who has already said that they can move them if they want to".

We seem to be arguing separate points -- I'm saying that it's not unreasonable that a small company could be moved fairly easily. Possibly as easily as unplugging a blade enclosure and throwing it in a station wagon. There are loads of small businesses that can run on an amount of hardware that can be easily transported. (When I used to gig on electric bass, my amp and other rack gear was in a portable 8U rack and that was more "portable" than the 100 pound speaker cabinet.)

I can't tell if your point is that it's unreasonable for all companies, which is wrong, or that it's unreasonable for some companies, which is obvious.

Re: Fog Creek is about to go down

#77
post #49

Earlier quoted context omitted.

Nobody was really expecting this much of Manhattan to lose power.

I'm sorry, but these chaps have been around the block a few times, they're not new start-ups. They know that there is no way on earth you can guarantee (or even reasonably be sure that) a single datacenter won't fail, even under non-emergency conditions, so their customers (I'm not one) should be calling them out on why they said that all would be A-OK. It would have been much better to say something like "We have pu…

You can be reasonably sure without being certain. They didn't guarantee anything. If you misread anything they've written as "guarantee" their services will never go down, you're blaming the wrong person. (Full disclosure: I am one of their customers)

Re: Fog Creek is about to go down

#78
post #69

Earlier quoted context omitted.

I meant ordinary citizens, not emergency planners. Despite being told that they could be without power, many of my friends did not believe it. "How could Manhattan lose power?"

> "How could Manhattan lose power?" Seriously? Were they not living in New York in 2003?

They were not.

I'm guessing everyone is basing their experience on last year's "hurricane", which was not nearly as bad as this one.

Re: Fog Creek is about to go down

#79
post #38
post #20

Earlier quoted context omitted.

Ofcourse they have redundancy, just not cross-datacenter redundancy. And if you knew anything about cross-datacenter redundancy you'd know that cross-datacenter redundancy is something you do not decide upon lightly. Then again, having cross-datacenter backups that can easily be taken online would be a bit more professional than 'we want to physically move the servers'.

I'll be the first to admit I don't really know anything about cross-datacenter redundancy; however, I always thought that was pretty high on the list once you had SaaS products that were pulling in enough revenue to warrant full-time employees outside of the founders. What are the reasons why you would choose not to do it? Are they all financial or are there other implications?

I think the biggest argument against complex cross-DC redundancy is that it can add complexity and failure modes, not just during the emergency, but every day.

As a simple example, I've seen at least a half dozen people who had issues because they thought it was as simple as throwing a mysql node into each datacenter, only to discover (much later) that the databases had become inconsistent and that failing over created bigger problems than it solved.

Similarly, I've seen complex high-availability infrastructures where the complexity of that infrastructure created more net downtime than a simpler infrastructure would've, it just went down at slightly different times.

And you really need to think about the implications of various failure modes. If you go down in the middle of a transaction, is that a problem for your application? Is it okay to roll back to data that's 3 hours old? 3 minutes? 3 seconds?

There are any number of situations where it's reasonable to say "we expect our datacenter will fail once every couple decades and when it does, we'll be down for a couple days."

Re: Fog Creek is about to go down

#80
post #41

Well, good luck to them, both personally and in bringing it back soon. I only wish Trello hadn't tried to reload on its own, so I could still see the screen before the shutdown. Now all I have is a blank page :(

I've managed to successfully failover to our own emergency backup version: https://twitter.com/williamlannen/status/263294924382937090
Post reply on HN