Live data from Hacker News

DataDog is having a major outage across almost all services

status.datadoghq.com

1–10 of 40 posts

Re: DataDog is having a major outage across almost all services

#4
All of Datadog's auto-muting logic during incidents is super well thought out and impresses me every time.

I'd have a few incidents open at 2am for missing business metrics and hosts falling out of the sky due to this if they didn't have that logic there, but instead we've sent out no false alerts for this.

Re: DataDog is having a major outage across almost all services

#6
post #3
post #2

It's been hard down for 5 hours now.

At least their self monitoring works.

I joked about this with a coworker, but I do have to wonder what they actually use for monitoring internally. It would be interesting if they just have a second copy of the prod stack for internal monitoring, or something.

Re: DataDog is having a major outage across almost all services

#7
post #6
post #3

Earlier quoted context omitted.

At least their self monitoring works.

I joked about this with a coworker, but I do have to wonder what they actually use for monitoring internally. It would be interesting if they just have a second copy of the prod stack for internal monitoring, or something.

We definitely used datadog to monitor internal systems while I was there many years ago. Dogfooding was a celebrated practice and I'd be surprised to learn that it's changed.

Re: DataDog is having a major outage across almost all services

#8

This effected a production deploy today - our team did not know datadog was down prior to deployment. Horrible.

We have "datadog outage" as an abort condition that is checked before any deploys or operations. I highly suggest implementing something like that for things you depend on.

If y'all are slinging code you are aware that nothing is 100% available, a consumer failing to anticipate that isn't really the providers fault.

Re: DataDog is having a major outage across almost all services

#9
post #7
post #6

Earlier quoted context omitted.

I joked about this with a coworker, but I do have to wonder what they actually use for monitoring internally. It would be interesting if they just have a second copy of the prod stack for internal monitoring, or something.

We definitely used datadog to monitor internal systems while I was there many years ago. Dogfooding was a celebrated practice and I'd be surprised to learn that it's changed.

I have to assume it's a separate stack from what is used in production?

Re: DataDog is having a major outage across almost all services

#10
post #4

All of Datadog's auto-muting logic during incidents is super well thought out and impresses me every time. I'd have a few incidents open at 2am for missing business metrics and hosts falling out of the sky due to this if they didn't have that logic there, but instead we've sent out no false alerts for this.

For those who haven’t used Datadog, what do you mean by auto-muting logic? What does it do exactly?
Post reply on HN