Live data from Hacker News

Mandrill has been down for over 30 hours with no explanation

twitter.com

51–60 of 115 posts

Re: Mandrill has been down for over 30 hours with no explanation

#51
post #48
post #44

Earlier quoted context omitted.

I moved to sendgrid when they announced mandrill's merger and have never been happier.

Except now, sendgrid is merged with twilio; so... history repeating?

No. Twilio has a status page :)

Re: Mandrill has been down for over 30 hours with no explanation

#52
post #33

Just remember that you can't trust their interface, even though their Outbound page says "Delivered" it isn't delivered unless there are SMTP events attached, if it looks like this, your email is queued, not sent: https://cdn.servnice.com/screenie/c1Og3TcrkLFV9Hg.jpg

exactly! this was the main reason I moved away.

Re: Mandrill has been down for over 30 hours with no explanation

#53
post #27

Unfortunately this is standard MailChimp way of doing things ever since they screwed over paying customers and merged with Mandrill [1]. They are the opposite of a transparent organization - they are the GoDaddy of this business. (I was super pissed once they removed their status page..WTF). We hit these errors a few months ago with our clients and I put up a roadmap to move all of them off the terrible joke that Man…

I really do hate posts like this.

if youre not going to name where you moved your customers too, then who cares?

Re: Mandrill has been down for over 30 hours with no explanation

#54
post #50
post #27

Unfortunately this is standard MailChimp way of doing things ever since they screwed over paying customers and merged with Mandrill [1]. They are the opposite of a transparent organization - they are the GoDaddy of this business. (I was super pissed once they removed their status page..WTF). We hit these errors a few months ago with our clients and I put up a roadmap to move all of them off the terrible joke that Man…

Where did you move them to?

I built a Mandrill clone based on SES, Google AppEngine, Phoenix/Elixir. I moved everything there and it works pretty well, planning to open source it at some point.

Re: Mandrill has been down for over 30 hours with no explanation

#55
post #53
post #27

Unfortunately this is standard MailChimp way of doing things ever since they screwed over paying customers and merged with Mandrill [1]. They are the opposite of a transparent organization - they are the GoDaddy of this business. (I was super pissed once they removed their status page..WTF). We hit these errors a few months ago with our clients and I put up a roadmap to move all of them off the terrible joke that Man…

I really do hate posts like this. if youre not going to name where you moved your customers too, then who cares?

Thanks for the feedback, I'll edit my post.

Re: Mandrill has been down for over 30 hours with no explanation

#58
Got this email just now.

- - - -

Hello,

We’re contacting you about an ongoing outage with the Mandrill app. This email provides background on what happened and how users are affected, what we’re doing to address the issue, and what’s next for our customers.

What happened Mandrill uses a sharded Postgres setup as one of our main datastores. On Sunday, February 3, at 10:30pm EST, 1 of our 5 physical Postgres instances saw a significant spike in writes. The spike in writes triggered a Transaction ID Wraparound issue. When this occurs, database activity is completely halted. The database sets itself in read-only mode until offline maintenance (known as vacuuming) can occur.

The database is large—running the vacuum process takes a significant amount of time and resources, and there’s no clear way to track progress.

Customer impact The impact to users could come in the form of not tracking opens, clicks, bounces, email sends, inbound email, webhook events, and more. Right now, it looks like the database outage is affecting up to 20% of our outbound volume as well as a majority of inbound email and webhooks.

What we’re doing to address this We don’t have an estimated time for when the vacuum process and cleanup work will be complete. While we have a parallel set of tasks going to try to get the database back in working order, these efforts are also slow and difficult with a database of this size. We’re trying everything we can to finish this process as quickly as possible, but this could take several days, or longer. We hope to have more information and a timeline for resolution soon.

In the meantime, it’s possible that you may see errors related to sending and receiving emails. We’ll continue to update you on our progress by email and let you know as soon as these issues are fully resolved.

What’s next We apologize for the disruption to your business. Once the outage is resolved, we plan to offer refunds to all affected users. You don’t need to take any action at this time—we’ll share details in a follow-up email and will automatically credit your account.

Again, we’re sorry for the interruption and we hope to have good news to share soon.

Re: Mandrill has been down for over 30 hours with no explanation

#59

Got this email just now. - - - - Hello, We’re contacting you about an ongoing outage with the Mandrill app. This email provides background on what happened and how users are affected, what we’re doing to address the issue, and what’s next for our customers. What happened Mandrill uses a sharded Postgres setup as one of our main datastores. On Sunday, February 3, at 10:30pm EST, 1 of our 5 physical Postgres instances…

Ha! For once it’s not MongoDB but Postgres. I wonder why the sending is effected though. Can’t they run their service with an empty databse in the meantime?

Re: Mandrill has been down for over 30 hours with no explanation

#60

Got this email just now. - - - - Hello, We’re contacting you about an ongoing outage with the Mandrill app. This email provides background on what happened and how users are affected, what we’re doing to address the issue, and what’s next for our customers. What happened Mandrill uses a sharded Postgres setup as one of our main datastores. On Sunday, February 3, at 10:30pm EST, 1 of our 5 physical Postgres instances…

Why does one shard being down remove all their inbound functionality? I'm struggling to understand the purpose of sharding if you can't pull a node offline and replace it while you deal with the wraparound. Is it part of postgres that if one shard has an issue, the entire cluster goes into read only mode?
Post reply on HN