DataDog is having a major outage across almost all services
status.datadoghq.com
DataDog is having a major outage across almost all services
1–10 of 40 posts
Re: DataDog is having a major outage across almost all services
#2Re: DataDog is having a major outage across almost all services
#3It's been hard down for 5 hours now.
Re: DataDog is having a major outage across almost all services
#4I'd have a few incidents open at 2am for missing business metrics and hosts falling out of the sky due to this if they didn't have that logic there, but instead we've sent out no false alerts for this.
Re: DataDog is having a major outage across almost all services
#5Re: DataDog is having a major outage across almost all services
#6It's been hard down for 5 hours now.
At least their self monitoring works.
Re: DataDog is having a major outage across almost all services
#7Earlier quoted context omitted.
At least their self monitoring works.
I joked about this with a coworker, but I do have to wonder what they actually use for monitoring internally. It would be interesting if they just have a second copy of the prod stack for internal monitoring, or something.
Re: DataDog is having a major outage across almost all services
#8This effected a production deploy today - our team did not know datadog was down prior to deployment. Horrible.
If y'all are slinging code you are aware that nothing is 100% available, a consumer failing to anticipate that isn't really the providers fault.
Re: DataDog is having a major outage across almost all services
#9Earlier quoted context omitted.
I joked about this with a coworker, but I do have to wonder what they actually use for monitoring internally. It would be interesting if they just have a second copy of the prod stack for internal monitoring, or something.
We definitely used datadog to monitor internal systems while I was there many years ago. Dogfooding was a celebrated practice and I'd be surprised to learn that it's changed.
Re: DataDog is having a major outage across almost all services
#10All of Datadog's auto-muting logic during incidents is super well thought out and impresses me every time. I'd have a few incidents open at 2am for missing business metrics and hosts falling out of the sky due to this if they didn't have that logic there, but instead we've sent out no false alerts for this.