Live data from Hacker News

Mandrill has been down for over 30 hours with no explanation

twitter.com

111–115 of 115 posts

Re: Mandrill has been down for over 30 hours with no explanation

#111
post #68

Got this email just now. - - - - Hello, We’re contacting you about an ongoing outage with the Mandrill app. This email provides background on what happened and how users are affected, what we’re doing to address the issue, and what’s next for our customers. What happened Mandrill uses a sharded Postgres setup as one of our main datastores. On Sunday, February 3, at 10:30pm EST, 1 of our 5 physical Postgres instances…

If you care about scalability and availability simultaneously, I'm not sure in these modern times why you would use a relational database. When they fail, they fail catastrophically and are difficult to recover, as this failure event (and the never-ending stream of failure events posted to HN) demonstrates. Don't get me wrong--I love relational databases and they are amazing pieces of technology. But they are incredi…

Transaction wrap around is very well known issue and easy to avoid with autovacumming.

Relational databases are tried and true and we have learned from the failures and have only made the technology better.

There are many use cases from data modeling perspective where a relational db makes more sense than a no sql and you really have to understand the trade offs of consistency and durability too. There will always be a place for both technologies and its not a question of either/or but rather what makes sense for your application in terms of not only system scalability but data scalability.

Re: Mandrill has been down for over 30 hours with no explanation

#112

Earlier quoted context omitted.

> They are also known to remove accounts on the system based on these partisan politics. Please provide proof when making claims like this. Infowars/Alex Jones doesn't count, they were universally blacklisted.

ROFL. Please provide proof they remove accounts based on partisan politics! Except for that case that they removed an account based on partisan politics!

[deleted]

Re: Mandrill has been down for over 30 hours with no explanation

#113
post #68

Got this email just now. - - - - Hello, We’re contacting you about an ongoing outage with the Mandrill app. This email provides background on what happened and how users are affected, what we’re doing to address the issue, and what’s next for our customers. What happened Mandrill uses a sharded Postgres setup as one of our main datastores. On Sunday, February 3, at 10:30pm EST, 1 of our 5 physical Postgres instances…

If you care about scalability and availability simultaneously, I'm not sure in these modern times why you would use a relational database. When they fail, they fail catastrophically and are difficult to recover, as this failure event (and the never-ending stream of failure events posted to HN) demonstrates. Don't get me wrong--I love relational databases and they are amazing pieces of technology. But they are incredi…

[deleted]

Re: Mandrill has been down for over 30 hours with no explanation

#114
post #68

Earlier quoted context omitted.

If you care about scalability and availability simultaneously, I'm not sure in these modern times why you would use a relational database. When they fail, they fail catastrophically and are difficult to recover, as this failure event (and the never-ending stream of failure events posted to HN) demonstrates. Don't get me wrong--I love relational databases and they are amazing pieces of technology. But they are incredi…

https://aws.amazon.com/message/5467D2/

What is your point? It was a 6 hour brownout, not a 30+ hour blackout. It is very unlikely that this kind of outage will happen again for DynamoDB. How likely is someone else going to run into a transaction wrap around again? If it's such a well-known issue, then presumably it keeps happening to a lot of people.

Re: Mandrill has been down for over 30 hours with no explanation

#115
post #68

Earlier quoted context omitted.

If you care about scalability and availability simultaneously, I'm not sure in these modern times why you would use a relational database. When they fail, they fail catastrophically and are difficult to recover, as this failure event (and the never-ending stream of failure events posted to HN) demonstrates. Don't get me wrong--I love relational databases and they are amazing pieces of technology. But they are incredi…

Transaction wrap around is very well known issue and easy to avoid with autovacumming. Relational databases are tried and true and we have learned from the failures and have only made the technology better. There are many use cases from data modeling perspective where a relational db makes more sense than a no sql and you really have to understand the trade offs of consistency and durability too. There will always be…

I'm not saying that you should never use relational databases. But if you are running at a large scale and have tight availability SLAs...then consider not using relational databases.

The fact that transaction wrap around is so well-known is itself a red flag--apparently a lot of people have run into this issue, and yet it keeps being an issue. The blast radius is very large and the recovery is painful, as shown here by Mandrill. You should think twice before accepting that risk if you value your uptime.

If you want to become an expert on all these pitfalls and caveats of running relational databases at scale, at the expense of your availability and customer satisfaction--then by all means continue using relational databases. For many use cases, there are better options with better failure resiliency and recovery stories.

Post reply on HN