Live data from Hacker News

DataDog is having a major outage across almost all services

status.datadoghq.com

11–20 of 40 posts

Re: DataDog is having a major outage across almost all services

#11
post #4

All of Datadog's auto-muting logic during incidents is super well thought out and impresses me every time. I'd have a few incidents open at 2am for missing business metrics and hosts falling out of the sky due to this if they didn't have that logic there, but instead we've sent out no false alerts for this.

For those who haven’t used Datadog, what do you mean by auto-muting logic? What does it do exactly?

We have alerts set up that expect metrics for things like "orders placed" to always be happening at expected rates.

When datadog has the very rare outage that breaks ingestion, all of our alerts would normally go off because we aren't seeing the expected volume of "orders placed" and open up StatusPage incidents for us and our customers, call the pagers and get folks working.

But instead they automatically stop any false alerts that would normally alert here because of their outage. Saves me a lot of headaches.

It is stuff like this why I am happy paying the Datadog bills. Even their outages are good.

Re: DataDog is having a major outage across almost all services

#13

We're up to 10 hours of downtime on their services. Anyone know what could have caused this? Companies generally [citation needed] don't go down for half a day across all their services.

I think having very infrequent large long outages is the norm here actually.

I can't think of any company this size that didn't have some outage of this magnitude at least once.

Facebook BGP for instance, Slack in Feb of '22, Cloudflare in June, YouTube, Twitch, Sony PlayStation etc etc have all had incidents this wide and long.

Re: DataDog is having a major outage across almost all services

#14
Ironically, DD promotes using their tool to set and measure SLAs but has a low bar on their SLA:

> Excluding scheduled maintenance windows, Datadog will use commercially reasonable efforts to maintain 99.8% availability of the hosted portion of the Service for each calendar month during the term of this Agreement. The Service will be deemed “available” so long as Authorized Users are able to login to the Service interface and access monitoring data. Excluding planned maintenance periods, in the event the Service availability drops below 99.8% for two consecutive months, Customer may terminate the Service in the calendar month following such two-month period upon written notice to Datadog. To assess uptime, Customer may, if under a Paying Plan, request the Service availability for a prior month by filing a support ticket through the Site.

Re: DataDog is having a major outage across almost all services

#15
post #11

Earlier quoted context omitted.

For those who haven’t used Datadog, what do you mean by auto-muting logic? What does it do exactly?

We have alerts set up that expect metrics for things like "orders placed" to always be happening at expected rates. When datadog has the very rare outage that breaks ingestion, all of our alerts would normally go off because we aren't seeing the expected volume of "orders placed" and open up StatusPage incidents for us and our customers, call the pagers and get folks working. But instead they automatically stop any f…

This is convenient behavior up until you actually have an incident that coincides with theirs, in which case it becomes catastrophic because you had no idea that outside vigilance was required on account of their ingestion downtime. Not sure why you would laud this. Is it possible to opt out?

Re: DataDog is having a major outage across almost all services

#16
post #14

Ironically, DD promotes using their tool to set and measure SLAs but has a low bar on their SLA: > Excluding scheduled maintenance windows, Datadog will use commercially reasonable efforts to maintain 99.8% availability of the hosted portion of the Service for each calendar month during the term of this Agreement. The Service will be deemed “available” so long as Authorized Users are able to login to the Service inte…

I don't understand how that is ironic.

Doesn't seem like that SLA could be defined as a "low bar" to me honestly, 99.8% in writing is impressive. It's public as well, meaning if you need a better SLA they aren't the ones for you.

New Relic is 98.5% https://docs.newrelic.com/docs/licenses/license-information/...

Re: DataDog is having a major outage across almost all services

#17
post #11

Earlier quoted context omitted.

We have alerts set up that expect metrics for things like "orders placed" to always be happening at expected rates. When datadog has the very rare outage that breaks ingestion, all of our alerts would normally go off because we aren't seeing the expected volume of "orders placed" and open up StatusPage incidents for us and our customers, call the pagers and get folks working. But instead they automatically stop any f…

This is convenient behavior up until you actually have an incident that coincides with theirs, in which case it becomes catastrophic because you had no idea that outside vigilance was required on account of their ingestion downtime. Not sure why you would laud this. Is it possible to opt out?

In your scenario you would have no logs etc until the DD incident resolved.

Opting out would just mean all your missing data alerts fire every time Datadog has an incident and you would then check, see that everything is missing, and then identify the cause as the Datadog incident.

Its much better to have them handle it and auto-mute the impacted monitors than communicate to my customers every time about false alerts saying all our services are down.

Re: DataDog is having a major outage across almost all services

#18
post #17

Earlier quoted context omitted.

This is convenient behavior up until you actually have an incident that coincides with theirs, in which case it becomes catastrophic because you had no idea that outside vigilance was required on account of their ingestion downtime. Not sure why you would laud this. Is it possible to opt out?

In your scenario you would have no logs etc until the DD incident resolved. Opting out would just mean all your missing data alerts fire every time Datadog has an incident and you would then check, see that everything is missing, and then identify the cause as the Datadog incident. Its much better to have them handle it and auto-mute the impacted monitors than communicate to my customers every time about false alerts…

> Opting out would just mean all your missing data alerts fire every time Datadog has an incident and you would then check, see that everything is missing, and then identify the cause as the Datadog incident.

You are missing the last step, which is that, knowing alerts are down, you can actively monitor using other tools/reporting for the duration of their incident.

And why would you have no logs? Even assuming you ingest logs through Datadog (they monitor on much than just logs and not everyone uses all facets of their offering), you would presumably have some way to access them more directly (even tailing output directly if necessary).

And lastly, why would you communicate to your customers without any idea of the scope or cause of the issue? It would likely be clear very quickly that Datadog was having issues when you see that all your metrics are suddenly discontinued without other ill effect.

Re: DataDog is having a major outage across almost all services

#19

This effected a production deploy today - our team did not know datadog was down prior to deployment. Horrible.

Out of curiosity, what datadog services do you use that cause deployments to be interrupted when they are having an outage?

Re: DataDog is having a major outage across almost all services

#20
> Mar 08, 2023 - 13:14 EST > Update - We are continuing to make progress towards recovering all services. Data ingestion and monitor notifications remain delayed across all data types.

> Mar 08, 2023 - 12:29 EST > Update - We continue progress towards recovering all services. Data ingestion and monitor notifications remain delayed across all data types.

> Mar 08, 2023 - 11:46 EST > Update - We continue progress towards recovering all services. Data ingestion and monitor notifications remain delayed across all data types.

-----

I won't post all of it, but you get the picture. Datadog's status updates go on like this for 12 hours. This is a status update anti-pattern. These updates add no value, except maybe some small reassurance that Datadog hasn't forgotten they are down, and they continue to work on it.

But it's actually worse than no update at all, because now every customer needs to parse through a whole mess of these updates to try and figure out what's happening, and if you subscribe to the updates, you start getting spam with no new information every 45 minutes.

I know companies are in a bind here. They don't want to provide estimates they might miss, or adhoc engineering info without vetting. On the other hand, customers complain about a lack of communication if there are no updates. But spamming your status page like this is not a real update, it's a pretend update. Just something to point to when customers complain about a lack of communication, but ultimately still a lack of communication.

Post reply on HN