Live data from Hacker News

Ask HN: Ever worked with a service that can never be restarted?

news.ycombinator.com

101–110 of 201 posts

Re: Ask HN: Ever worked with a service that can never be restarted?

#101
post #42

Earlier quoted context omitted.

Everyone's afraid to do anything on that machine at all. So even getting agreement to login or install software or run tcpdump or whatever is fraught.

What I was meaning was that if you can get a hold of the binary that is running you could take it to some other environment and have a look (starting with a hex dump and working up from there...). Mind you if things are that bad then maybe I'd follow my own advice (in another comment) and stay well away from the currently running thing.

this would certainly help recreate the configuration.

Re: Ask HN: Ever worked with a service that can never be restarted?

#102
post #58
post #8

Earlier quoted context omitted.

Oh for sure this is possible: https://www.youtube.com/watch?v=vQ5MA685ApE the devices that do it without doing what these guys did is quite pricey though https://www.cru-inc.com/products/wiebetech/hotplug_field_kit...

That requires the computer to hotswap to be plugged into a power strip. It doesn't work if the computer is plugged into the wall directly. (Well leaving aside modifying the house's electricity / opening up the wall obviously.)

Incorrect. See this video[0] from the product stage. They have two ways to hotswap if it's not plugged into a powerbar: 1) plugging in their device into the adjacent power outlet then removing the socket assembly and 2) an adapter that can splice leads into the power cable

https://youtu.be/-G8sEYCOv-o

Pretty ingenious, tbh. Though I'm a bit concerned about how much you end up handling energized plugs.

Re: Ask HN: Ever worked with a service that can never be restarted?

#103

Earlier quoted context omitted.

Sorry that nobody else got it, but I like that joke. Do you think we can put some crypto in there, too?

> Sorry that nobody else got it, but I like that joke. No, plenty of us got it. It's just not funny right now. A person is trying to find a real solution to a real problem they have. Every comment in this thread represents a potential hope that someone knows how to fix it. When someone comes in to a serious thread that is spitballing ideas and tries to make it about them by cracking jokes and being fun and clever, it…

what a silly excuse for being offended by nothing.

Watch as I come up with a random reason you're wrong.

The OP could see it and laugh and that could possibly lighten their day and help them relax.

Wait, no. That's positive, online we're only supposed to bash men, hitler, and take everything in the most negative light possible. I should have instead judged you as a hateful person who has a horrible life and no friends.

Re: Ask HN: Ever worked with a service that can never be restarted?

#104
post #88

I once worked with a client in such a situation. They were in the process of building a huge oil rig, costing somewhere north of $5.5bn USD. The basic premise was that a previous vendor had configured a big documentation system running on a soon-to-be-outdated Windows Server version ages ago, along with a "kind of" API allowing the shipyard to send information in the form of equipment/construction metadata, documents…

> In the end everything was kind of anti-climactic, with everything working as expected for as long as it needed to.

I often tell people that great software development is boring.

Re: Ask HN: Ever worked with a service that can never be restarted?

#105
post #31
post #20

Earlier quoted context omitted.

> I'm not sure that 55 months of up-time indicates it's more or less likely to go down in the next month, but I'd guess more likely. If 'going down' is an event governed by chance (e.g. power failure), then it does not matter if it has been up for 55 months or 55 minutes. https://en.wikipedia.org/wiki/Gambler%27s_fallacy

There’s also the Turkey fallacy - Turkeys that are smug about knowing the Gambler’s fallacy keep thinking their survival is independent of their age, until people eat them. My probability of dying goes up with each year as well, it’s not constant. Think the probability of the machine shutting down in the next minute is independent of age, but it shutting off in the next 100 years is certain. There’s a line or curve t…

This is how I feel when people talk about Schroedingers cat. Give me a week and I'll be able to answer that question definitively.

Re: Ask HN: Ever worked with a service that can never be restarted?

#106
post #97

My advice: 1. Suggest that the work to replace it is prioritized commensurate with the business impact caused whilst re-establishing service as if it went down right now 2. Remind them that it will go down at the worst possible time. 3. Ensure your name is attached to these two warnings. 4. Promise yourself you wouldn't run your business this way. 5. Get on with your life.

What verytrivial said!!!! I’ve been in similar situations, but never without an immediately obvious solution. That system WILL FAIL. Even as we speak the time-to-failure is shrinking. Even if one of the solutions described below actually works, you won’t get 100% recovery. I recall a story years ago - from MIT, if memory serves - where they rebooted a system because they had many generations of Sybase backups. When t…

Aah, my favorite write-only backups.

Re: Ask HN: Ever worked with a service that can never be restarted?

#107
post #98
post #92

Unpopular advice: find a subtle way to make it crash, preferably stealthy, but if not possible, at least in a way that can be attributable to mild innocent incompetence instead of malice! Then there will be more and more interesting work to do for you and others, either rediscovering and properly documenting the config, or, hopefully, architecting and coding its replacement! In the aftermath, the organization will be…

Back when I ran a company, my partners and I gamed out what to do with employees in various situations, as a way of coming to a consensus. I was against legal action in all but extreme hypotheticals. Behavior like this is in that class. You assume you not only have all the relevant information to predict the outcome, but also that your analysis of the situation is superior enough to trump those who actually own the p…

If you had hired competent people to begin with, the configs would have been backed up.

I don't agree with the other posters recommendation, but I understand the sentiment. The same incompetence that allowed the event to happen is the same incompetence that allowed it to be swept under the rug for almost 6 years.

I've seen paragons of hiring get upset because failures were made visible, not understanding that the failures were there all along and they were playing a game of chicken with reality, only now they have a chance of winning.

Re: Ask HN: Ever worked with a service that can never be restarted?

#108
In South Africa the citizens are facing a crises of such a service and its our electricity system. Predominately coal power plants, that are at end of life, and rolling blackouts is a common occurrence. It is necessary action, to avoid a blackout, given it would take weeks to get the system going again. Furthermore, high inequality and unemployment, makes it paramount to grow the economy, so these rolling blackouts have a huge impact socio-ecomically. It is a service that can't simply be restarted.

Transitioning risky services which can't be restarted, is clearly complicated, especially for complex systems. I wonder if there is body of knowledge with principles that could apply to not only to these massive nationwide scale, but to that of web-services and alike. Does anybody have resources in this respect?

Re: Ask HN: Ever worked with a service that can never be restarted?

#109
Interesting problem. There's some good ideas in here, and I think you might get some better ones if you add more details.

You said that you have been assigned the task of replacing the "zombie" program, which I assume means that you are to write a new one that interfaces whatever depends on the zombie.

If the zombie were to die right now, how severe would the consequences be? On the scale of total disruption of business to minor inconvenience that could be worked around until it's fixed?

How complex is the zombie? What is the nature of the state that was supplied by the lost configuration file? Is it stuff like which port to listen on and how to connect to its database, or is it more like huge tables of magic values that won't be easy to figure out again?

Do you currently have enough information about the zombie to write a replacement? Do you have what you need to test your replacement before deploying it?

If you are confident you have sufficient information to write the replacement, how long do you think you need to write and test that?

You said the zombie runs on a physical server. What operating system? What language/runtime/stack is the zombie based on?

Re: Ask HN: Ever worked with a service that can never be restarted?

#110
post #97

My advice: 1. Suggest that the work to replace it is prioritized commensurate with the business impact caused whilst re-establishing service as if it went down right now 2. Remind them that it will go down at the worst possible time. 3. Ensure your name is attached to these two warnings. 4. Promise yourself you wouldn't run your business this way. 5. Get on with your life.

What verytrivial said!!!! I’ve been in similar situations, but never without an immediately obvious solution. That system WILL FAIL. Even as we speak the time-to-failure is shrinking. Even if one of the solutions described below actually works, you won’t get 100% recovery. I recall a story years ago - from MIT, if memory serves - where they rebooted a system because they had many generations of Sybase backups. When t…

> That system WILL FAIL.

I think that the key point. We have a system that can't easily be restarted and won't automatically start after a server is rebooted. The developers of the software don't care, because "what are the chances of a virtual machine spontaneously restarting". Turns out, those chances are rather good.

Servers, virtual machines, containers, doesn't matter, unless that thing is running on a mainframe, it will crash.

Post reply on HN