I'll share how wavii does this. We use amazon elb and chef. The first chef recipe that runs on our frontend is "touch /tmp/website_should_drain". The website then returns 404 from the /status health check URL. We wait 30s, update the site, and then rm the draining sentinel file. If the deployment fails the host stays out of the elb. If more than 1/3 host fail we abort the whole deployment. This works well for us and…
We've found that when the ELB health check returns non-ok, the ELB will immediately kill any requests in flight to that box. So it's kind of impossible to do a real "drain" with ELB. Are you seeing different behavior?
How we handle deploys and failover without disrupting user experience
11–16 of 16 posts
Re: How we handle deploys and failover without disrupting user experience
#12By far the easiest way to get similar behavior is to run your Django app via Gunicorn (proxied from Nginx). Gunicorn supports hot code reloading via SIGHUP, and it does so by forking and gracefully killing old processes. If your requirements don't match gunicorn (not django, not python, etc) then you can use https://github.com/TimothyFitz/zdd a project I wrote to automate rewriting nginx config files to deal with cha…
I mentioned this in a thread the other day; I've yet to see a good use case for hot code reloading. Can you really not drain requests to that host via HAProxy (or similar), and then actually restart the service? The nice thing about that approach is that your choice of service runtime doesn't matter.
Zdd (the project I linked to) is all about spawning a new process in parallel. All the advantages of your approach (switch to an entirely different language? Who cares) but without the stalls.
Zdd also lets you keep the old process alive through the duration of the deploy (and after), and with a little work could let you switch back in the event of a bad deploy without having to start the old version up again.
Re: How we handle deploys and failover without disrupting user experience
#13Re: How we handle deploys and failover without disrupting user experience
#14Earlier quoted context omitted.
I mentioned this in a thread the other day; I've yet to see a good use case for hot code reloading. Can you really not drain requests to that host via HAProxy (or similar), and then actually restart the service? The nice thing about that approach is that your choice of service runtime doesn't matter.
Your method will stall responses for server shutdown + server startup time, which for Ruby/Python apps is usually measured in tens of seconds, and for other web servers can be much worse. Hot code reloading lets you avoid any downtime at all, and with it usually being built into the framework/language specific server you get the functionality "for free". Zdd (the project I linked to) is all about spawning a new proce…
The advantages to this method include being completely platform agnostic, as well as giving you a window to verify successful update without worrying about production traffic.
Re: How we handle deploys and failover without disrupting user experience
#15* Have more than one application server
* Deploy new code to a non-active application server
* Send some traffic to the new app server (perhaps based on cookie)
* When confident, switch traffic (using Varnish/other load balancer) to the new application server
Re: How we handle deploys and failover without disrupting user experience
#16Earlier quoted context omitted.
Your method will stall responses for server shutdown + server startup time, which for Ruby/Python apps is usually measured in tens of seconds, and for other web servers can be much worse. Hot code reloading lets you avoid any downtime at all, and with it usually being built into the framework/language specific server you get the functionality "for free". Zdd (the project I linked to) is all about spawning a new proce…
The method he's describing requires no downtime or stalled requests. Connections are drained from some pool in a load balance set if servers. The services is restarted with the new code. The hosts are given traffic again once the are initialized and healthy. The advantages to this method include being completely platform agnostic, as well as giving you a window to verify successful update without worrying about produ…
Personally I've found the "put new instances into a load balancer" method to make more sense for system changes (packages, kernels, OS versions) where deploying the change is inherently slow or expensive, but the method doesn't make sense for code deploys where deploy time is important.