Live data from Hacker News

Auth0 Down

twitter.com

31–40 of 78 posts

Re: Auth0 Down

#32
post #17

This is crazy timing -- my co-host and I just released a Podcast episode yesterday sifting through the details about the Auth0 database-related outage from 2018 ( https://downtimeproject.com/ ). I'll be curious to see how much overlap or not there is with that previous outage. They wrote up a nice post-mortem back then, so hopefully we'll get another one this time.

Do you have a link to that previous post-mortem?

https://cdn.auth0.com/blog/20181128-Incident-RCA.pdf

Re: Auth0 Down

#36
Wow, finding out just how many services use Auth0 because of this. Like I can't log into Segment (getting a 500).

I guess Okta did buy something core to the Saas ecosystem

Re: Auth0 Down

#37

We managed to get a reply from a C level. All we could get out of them was "something to do with our DB, but we don't know the root yet. Our fail-over process didn't work. This will never happen again". Also, it only took them 2 and a half hours to admit it was their entire system instead of "a small subset of users" lol.

Wasn't their previous major outage because of a bad migration?

Re: Auth0 Down

#38

What alternatives to Auth0 are worth looking into? Between this P0 (with no ability to check the status or file a ticket) and the Okta acquisition, I hesitate to continue using Auth0 as the default when spinning up new web apps.

AWS Cognito works but it is a far cry from how usable Auth0 is.

My blood pressure is still coming back down from AWS having an all-day outage during a holiday week last year, because Kinesis had an issue and it turns out their entire infrastructure depends on that.

Re: Auth0 Down

#39
post #37

We managed to get a reply from a C level. All we could get out of them was "something to do with our DB, but we don't know the root yet. Our fail-over process didn't work. This will never happen again". Also, it only took them 2 and a half hours to admit it was their entire system instead of "a small subset of users" lol.

Wasn't their previous major outage because of a bad migration?

I don't think so, I think that it was a combo of malicious intent and some indexes that never got run. I guess you might call it a bad migration since indexes didn't get run, but that seems more like a catalyst than a root. https://cdn.auth0.com/blog/20181128-Incident-RCA.pdf
Post reply on HN