Live data from Hacker News

How we handle deploys and failover without disrupting user experience

code.mixpanel.com

1–10 of 16 posts

Re: How we handle deploys and failover without disrupting user experience

#2
By far the easiest way to get similar behavior is to run your Django app via Gunicorn (proxied from Nginx). Gunicorn supports hot code reloading via SIGHUP, and it does so by forking and gracefully killing old processes.

If your requirements don't match gunicorn (not django, not python, etc) then you can use https://github.com/TimothyFitz/zdd a project I wrote to automate rewriting nginx config files to deal with changing proxied portfiles. To integrate any existing server, all you have to do is make it bind to port 0 (let the OS choose a port) and the write a foo.port file that contains the port number (like a pidfile). That's it.

Re: How we handle deploys and failover without disrupting user experience

#4
I'll share how wavii does this. We use amazon elb and chef. The first chef recipe that runs on our frontend is "touch /tmp/website_should_drain". The website then returns 404 from the /status health check URL. We wait 30s, update the site, and then rm the draining sentinel file. If the deployment fails the host stays out of the elb. If more than 1/3 host fail we abort the whole deployment. This works well for us and was very simple to implement.

Re: How we handle deploys and failover without disrupting user experience

#5
Mixpanel had some downtime 2 days ago, from 22:59 until 23:23 UTC according to our error logs. It might have only been for API users, but downtime nonetheless.

It's certainly possible it was scheduled, but it's not clear where that downtime is announced if so.

Re: How we handle deploys and failover without disrupting user experience

#8

By far the easiest way to get similar behavior is to run your Django app via Gunicorn (proxied from Nginx). Gunicorn supports hot code reloading via SIGHUP, and it does so by forking and gracefully killing old processes. If your requirements don't match gunicorn (not django, not python, etc) then you can use https://github.com/TimothyFitz/zdd a project I wrote to automate rewriting nginx config files to deal with cha…

Thanks for the Gunicorn hot-reloading tip. I've been using Gunicorn with Supervisor, and I wish I'd known about the SIGHUP trick earlier.

Re: How we handle deploys and failover without disrupting user experience

#9

I'll share how wavii does this. We use amazon elb and chef. The first chef recipe that runs on our frontend is "touch /tmp/website_should_drain". The website then returns 404 from the /status health check URL. We wait 30s, update the site, and then rm the draining sentinel file. If the deployment fails the host stays out of the elb. If more than 1/3 host fail we abort the whole deployment. This works well for us and…

We've found that when the ELB health check returns non-ok, the ELB will immediately kill any requests in flight to that box. So it's kind of impossible to do a real "drain" with ELB. Are you seeing different behavior?

Re: How we handle deploys and failover without disrupting user experience

#10

By far the easiest way to get similar behavior is to run your Django app via Gunicorn (proxied from Nginx). Gunicorn supports hot code reloading via SIGHUP, and it does so by forking and gracefully killing old processes. If your requirements don't match gunicorn (not django, not python, etc) then you can use https://github.com/TimothyFitz/zdd a project I wrote to automate rewriting nginx config files to deal with cha…

I mentioned this in a thread the other day; I've yet to see a good use case for hot code reloading. Can you really not drain requests to that host via HAProxy (or similar), and then actually restart the service? The nice thing about that approach is that your choice of service runtime doesn't matter.
Post reply on HN