Live data from Hacker News

Deno’s July 13th incident update

deno.com

11–20 of 47 posts

Re: Deno’s July 13th incident update

#11
> several services provided by the Deno company experienced a service disruptions in our us-west3 region for a period of just over 24 hours.

'"JUST" over 24 hours', no big deal of course /s

https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...

no mention of that issue

either nobody uses Deno so 0 complains

or people use Deno and for some reasons 24h+ downtime didn't impact anybody, wich is surprising, to say the least

Re: Deno’s July 13th incident update

#12

> several services provided by the Deno company experienced a service disruptions in our us-west3 region for a period of just over 24 hours. '"JUST" over 24 hours', no big deal of course /s https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que ... no mention of that issue either nobody uses Deno so 0 complains or people use Deno and for some reasons 24h+ downtime didn't impact anybody, wich is surprising, to s…

“Just over” 24 hours. As in slightly more than 24 hours. Which is different than “just” 24 hours.

Re: Deno’s July 13th incident update

#13

> several services provided by the Deno company experienced a service disruptions in our us-west3 region for a period of just over 24 hours. '"JUST" over 24 hours', no big deal of course /s https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que ... no mention of that issue either nobody uses Deno so 0 complains or people use Deno and for some reasons 24h+ downtime didn't impact anybody, wich is surprising, to s…

“Just over” 24 hours. As in slightly more than 24 hours. Which is different than “just” 24 hours.

> On July 13th, at around 18:45 UTC we started to receive reports of an outage

> On July 14th, at 19:14 UTC we were able to identify that the problem was within our us-west3 region

So yeah, 24 hours and 29 minutes.

Re: Deno’s July 13th incident update

#14
Is anyone running mission critical software on this rather new platform? In my experience every added cloud service becomes another potential weak link in your chain. A distributed DB for your data, a CDN for static assets, a couple of lambda functions for background processing. With every move away from the monolith your surface for potential downtime or "elevated error rates" increases.

Re: Deno’s July 13th incident update

#15

Hugops to the team. A quick question: is it intentional that there's nothing on https://denostatus.com/ ?

What would be the benefit for Deno to have an accurate status page? It would only stand to detract potential customers/investors

You guys took my comment way too literally. It was tongue in cheek :P

Re: Deno’s July 13th incident update

#16
> On July 13th, at around 18:45 UTC we started to receive reports of an outage from a small number of users. We investigated the status of our services, but were unable to confirm any of the reports. All of our status monitoring and tests reported that everything was operating normally.

> Over the course of the outage, we continued to monitor our service status, and worked with some of the affected users to narrow down the source of the problem.

> On July 14th, at 19:14 UTC we were able to identify that the problem was within our us-west3 region, which we then took offline, directing traffic to other nearby regions instead.

The time difference between when the first reports came in and when it was confirmed is a little concerning.

As an aside:

> ... approximately 18:00 UTC ...

> ... just over 24 hours ...

> ... For a period of around 24 hours, some users in the us-west3 region

> ... less than 30 minutes ...

> ... On July 13th, at around 18:45 UTC we started to receive reports of an outage from a small number of users. ...

"Approximately", "just", "around", "some", "small number of". It goes on and on. I disagree with the stylistic approach of being less specific in posts like these. A "small number of users" is relative. As readers, we have no idea what your typical load may be. Small may be a large number to us. "Just" over 24 hours is 26 hours? 24.5 hours? I implore you to be specific when you have the actual data.

These terms read as weasel words, and impact your effort at being fully transparent.

Re: Deno’s July 13th incident update

#17
I don't mean anything bad to Deno's team (I'm very partial to what they're building), but I'm rather surprised whenever a widely-publicized service has an outage that lasts hours or more than 24h. I'm genuinely curious to understand whether it's typically due to complexity of infrastructure and how hard it is to find route causes, how long it takes to redirect traffic / patch temporarily when the cause is found, or is it due to attitude where it's considered normal for these things to happen, and to take time to solve step by step.

Our services are of what I consider medium complexity (~70 services, ~10 different "layers" of logic, db, caching, load balancing etc, AWS, mostly self-managed centralized logging and monitoring) but still quite low-volume (We're very modestly funded compared to Deno (in this example) and the team is small...

Not sure whether that changes with traffic volume, complexity, team size, or is more primarily attitude-based and should continue to be cultivated.

Re: Deno’s July 13th incident update

#18
post #17

I don't mean anything bad to Deno's team (I'm very partial to what they're building), but I'm rather surprised whenever a widely-publicized service has an outage that lasts hours or more than 24h. I'm genuinely curious to understand whether it's typically due to complexity of infrastructure and how hard it is to find route causes, how long it takes to redirect traffic / patch temporarily when the cause is found, or i…

> is spontaneously met by my team as absolute emergency and typically fixed in Unless they had the "route the whole region over another one" in their prepared and practiced DR procedure, it would take any team a significant time to get that planned, approved, implemented and tested.

If you're running something at tens of services scale and recovered in 10min, you're extremely lucky. I'd suggest that if you don't have risks on your list that will take hours to resolve, your list is not complete.

Re: Deno’s July 13th incident update

#19
post #3

Ew, we've had similar issues in the past. These are really messy and confusing to recognize. In our case, 1 out of 5 LB instances lost its connection to the service discovery and later on ended up not knowing about a failover of one of the 5 backends for a service. As a result, something like 1 in 20 to 1 in 25 requests got answered with a connection refused. That took a minute to find.

Had something similar when a k8s node broke but k8s thought the pods (envoy) on it were still running so it routed 1/nth of traffic into a black hole

Re: Deno’s July 13th incident update

#20
post #16

> On July 13th, at around 18:45 UTC we started to receive reports of an outage from a small number of users. We investigated the status of our services, but were unable to confirm any of the reports. All of our status monitoring and tests reported that everything was operating normally. > Over the course of the outage, we continued to monitor our service status, and worked with some of the affected users to narrow do…

[deleted]
Post reply on HN