Live data from Hacker News

Mandrill has been down for over 30 hours with no explanation

twitter.com

61–70 of 115 posts

Re: Mandrill has been down for over 30 hours with no explanation

#61

Got this email just now. - - - - Hello, We’re contacting you about an ongoing outage with the Mandrill app. This email provides background on what happened and how users are affected, what we’re doing to address the issue, and what’s next for our customers. What happened Mandrill uses a sharded Postgres setup as one of our main datastores. On Sunday, February 3, at 10:30pm EST, 1 of our 5 physical Postgres instances…

If you're looking for a good series of blog posts about xid wraparound in Postgres check out these posts by Josh Berkus:

http://www.databasesoup.com/2012/09/freezing-your-tuples-off...

http://www.databasesoup.com/2012/10/freezing-your-tuples-off...

http://www.databasesoup.com/2012/12/freezing-your-tuples-off...

And this more recent one by Robert Haas:

https://rhaas.blogspot.com/2018/01/the-state-of-vacuum.html

As Josh states at the end of the third post the current best practices for dealing with this are really workarounds and as Robert states it requires monitoring and management. Postgres is an amazing piece of software and managing this is doable but IMHO this is one of Postgres' worst warts. It would be awesome if someone could donate some funding to improve this.

Re: Mandrill has been down for over 30 hours with no explanation

#62

Got this email just now. - - - - Hello, We’re contacting you about an ongoing outage with the Mandrill app. This email provides background on what happened and how users are affected, what we’re doing to address the issue, and what’s next for our customers. What happened Mandrill uses a sharded Postgres setup as one of our main datastores. On Sunday, February 3, at 10:30pm EST, 1 of our 5 physical Postgres instances…

Its also shared here: https://mailchi.mp/953cb079837a/important-information-about-...

Re: Mandrill has been down for over 30 hours with no explanation

#63
post #54
post #50

Earlier quoted context omitted.

Where did you move them to?

I built a Mandrill clone based on SES, Google AppEngine, Phoenix/Elixir. I moved everything there and it works pretty well, planning to open source it at some point.

How about you just open source it as a blog post and let someone else's clean it up. Heck even blogging the architecture and the issues would help.

Also are you using Amazon ses with Google app engine? Seems odd.

Re: Mandrill has been down for over 30 hours with no explanation

#64
post #49
post #27

Unfortunately this is standard MailChimp way of doing things ever since they screwed over paying customers and merged with Mandrill [1]. They are the opposite of a transparent organization - they are the GoDaddy of this business. (I was super pissed once they removed their status page..WTF). We hit these errors a few months ago with our clients and I put up a roadmap to move all of them off the terrible joke that Man…

yeah i just noticed the status page was removed - it used to actually be pretty good!

What an odd trend. If anything every company is clamoring to get a status page, not remove the one they have.

Re: Mandrill has been down for over 30 hours with no explanation

#65
I live in Atlanta where MailChimp (the parent company) is headquartered. They hire for progressive politics first, actual technical skills second/third. Many of their employees spend a large part of their day on social justice crusades targeting community members for wrong-think instead of gasp actually building a better products and taking care of customers. They are also known to remove accounts on the system based on these partisan politics.

At some point this is bound to catch up to them as competitors actually invest in their product. In the meantime, they will try to advertise the product like crazy and enough people will buy it without actually researching their actual customer service until this void is filled etc..

Re: Mandrill has been down for over 30 hours with no explanation

#66
post #54
post #50

Earlier quoted context omitted.

Where did you move them to?

I built a Mandrill clone based on SES, Google AppEngine, Phoenix/Elixir. I moved everything there and it works pretty well, planning to open source it at some point.

That sounds awesome! Please post a show hn when you do.

Re: Mandrill has been down for over 30 hours with no explanation

#67

Got this email just now. - - - - Hello, We’re contacting you about an ongoing outage with the Mandrill app. This email provides background on what happened and how users are affected, what we’re doing to address the issue, and what’s next for our customers. What happened Mandrill uses a sharded Postgres setup as one of our main datastores. On Sunday, February 3, at 10:30pm EST, 1 of our 5 physical Postgres instances…

If you're looking for a good series of blog posts about xid wraparound in Postgres check out these posts by Josh Berkus: http://www.databasesoup.com/2012/09/freezing-your-tuples-off... http://www.databasesoup.com/2012/10/freezing-your-tuples-off... http://www.databasesoup.com/2012/12/freezing-your-tuples-off... And this more recent one by Robert Haas: https://rhaas.blogspot.com/2018/01/the-state-of-vacuum.html As Jos…

My admittedly very superficial understanding of this issue is that the most common way to run into the xid wraparound problem is tuning the autovacuum in the wrong direction. So you notice that vacuum is taking up a lot of your servers resources, and decrease the frequency. Or you notice that it can't really keep up, but don't tune it to be more aggressive or provide enough resources for it to do its job. Or you don't monitor this at all, which is a pretty bad idea if you do billions of transactions (with less you can't really hit this issue).

This is also a problem that gets far harder to fix once you've run into it. If you have sufficient transaction volume to potentially hit this, you need to monitor autovacuum and make adjustments early before you get close to the wraparound. If you don't, you suddenly have to perform all the vacuum work at once, blocking that table until it's done.

Re: Mandrill has been down for over 30 hours with no explanation

#68

Got this email just now. - - - - Hello, We’re contacting you about an ongoing outage with the Mandrill app. This email provides background on what happened and how users are affected, what we’re doing to address the issue, and what’s next for our customers. What happened Mandrill uses a sharded Postgres setup as one of our main datastores. On Sunday, February 3, at 10:30pm EST, 1 of our 5 physical Postgres instances…

If you care about scalability and availability simultaneously, I'm not sure in these modern times why you would use a relational database. When they fail, they fail catastrophically and are difficult to recover, as this failure event (and the never-ending stream of failure events posted to HN) demonstrates.

Don't get me wrong--I love relational databases and they are amazing pieces of technology. But they are incredibly hard to "do right" at scale while maintaining availability SLAs.

edit:

I would appreciate if downvoters would explain their decision to downvote, so that if I'm incorrect then I could at least update my beliefs. My position is based on years of experience watching relational databases maintained by professional DBAs catastrophically fail in strange ways, and subsequently taking a long time to recover, causing complete blackouts. And having yet to see such failures in managed NoSQL DBs like DynamoDB.

Re: Mandrill has been down for over 30 hours with no explanation

#69
post #44
post #27

Unfortunately this is standard MailChimp way of doing things ever since they screwed over paying customers and merged with Mandrill [1]. They are the opposite of a transparent organization - they are the GoDaddy of this business. (I was super pissed once they removed their status page..WTF). We hit these errors a few months ago with our clients and I put up a roadmap to move all of them off the terrible joke that Man…

I moved to sendgrid when they announced mandrill's merger and have never been happier.

We moved to sendgrid around the same time. Our deliverability took a hit, and I'm not the keenest on their very narrow window of email data they keep... but no real shenanigans, very straight forward, and the one significant outage of theirs I can remember they handled it well.

Re: Mandrill has been down for over 30 hours with no explanation

#70
post #51
post #48

Earlier quoted context omitted.

Except now, sendgrid is merged with twilio; so... history repeating?

No. Twilio has a status page :)

And twilio seems to be very serious about uptime. Not sure if it translated to all the facets of their business, but phone services being down will lose you customers quickly.
Post reply on HN