Got this email just now. - - - - Hello, We’re contacting you about an ongoing outage with the Mandrill app. This email provides background on what happened and how users are affected, what we’re doing to address the issue, and what’s next for our customers. What happened Mandrill uses a sharded Postgres setup as one of our main datastores. On Sunday, February 3, at 10:30pm EST, 1 of our 5 physical Postgres instances…
http://www.databasesoup.com/2012/09/freezing-your-tuples-off...
http://www.databasesoup.com/2012/10/freezing-your-tuples-off...
http://www.databasesoup.com/2012/12/freezing-your-tuples-off...
And this more recent one by Robert Haas:
https://rhaas.blogspot.com/2018/01/the-state-of-vacuum.html
As Josh states at the end of the third post the current best practices for dealing with this are really workarounds and as Robert states it requires monitoring and management. Postgres is an amazing piece of software and managing this is doable but IMHO this is one of Postgres' worst warts. It would be awesome if someone could donate some funding to improve this.