An Update on Our Outage
21–30 of 235 posts
Re: An Update on Our Outage
#22Did they run out of internal ipv4 address space somehow? I’ll be curious for the post-mortem, this is definitely the longest outage of a global web service I can remember in… forever? This incident puts them under 2 9’s.
My guess is something related to service mesh/discovery. If that goes down, deployment infra also depends on it, and you don't have a solid black start protocol then you're in for a very bad week.
Re: An Update on Our Outage
#23It’s great to hear that it wasn’t the result of something malicious (e.g. a hack). But this has to be one of the longest outages by a company this big ($50B market cap), at least in the past decade it seems?
Anybody optimizing for competence is exchanging time for food and shelter and just gets to be reminded of the much harder game they are playing that produces much less optimal results of money.
Re: An Update on Our Outage
#24Earlier quoted context omitted.
What does that tend to look like? Do you have any examples?
In my experience, it tends to be that you find out one of your components has some tipping point as you scale out horizontally where everything goes to hell. Random example: you scale a service horizontally and suddenly postgres disk space usage goes from "this line looks horizontal" to "we just used 2 years' worth of disk use growth in 40 seconds". And the postmortem is basically like "yep can't really blame anyone…
Re: An Update on Our Outage
#25Re: An Update on Our Outage
#26On Thursday afternoon, October 28th, users began having trouble connecting with our platform. This immediately became our highest priority. Teams began working around the clock to identify the source of the problem and get things back to normal.
This was an especially difficult outage in that it involved a combination of several factors. A core system in our infrastructure became overwhelmed, prompted by a subtle bug in our backend service communications while under heavy load. This was not due to any peak in external traffic or any particular experience. Rather the failure was caused by the growth in the number of servers in our datacenters. The result was that most services at Roblox were unable to effectively communicate and deploy.
Due to the difficulty in diagnosing the actual bug, recovery took longer than any of us would have liked. Upon successfully identifying this root cause, we were able to resolve the issue through performance tuning, re-configuration, and scaling back of some load. We were able to fully restore service as of this afternoon.
We will publish a post-mortem with more details once we’ve completed our analysis, along with the actions we’ll be taking to avoid such issues in the future. In addition, we will implement a policy to make our creator community economically whole as a result of this outage. There are more details on this to come. As part of our “Respect the Community” value, we will continue to be transparent in our post-mortem.
To the best of our knowledge, there has been no loss of player persistence data, and your Roblox experience should now be fully back to normal. You can always contact our support team if you experience any hiccups using Roblox now or in the future.
We are grateful for the patience and support of our players, developers, and partners during this time.
Verbatim Text post for anyone having redirect issues.
TLDR; We have no clue what happened.
Re: An Update on Our Outage
#27Looks like everyone got the cause of this outage wrong on previous speculative posts.
Re: An Update on Our Outage
#28As most of the Roblox community is aware, we recently experienced an extended outage across our platform. We are sorry for the length of time it took us to restore service. A key value at Roblox is “Respect the Community,” and in this case, we apologize for the inconvenience to our community. On Thursday afternoon, October 28th, users began having trouble connecting with our platform. This immediately became our high…
Re: An Update on Our Outage
#29Can anyone cache this page and link that? Or paste the text here in a comment? This link shows a redirect loop for me.
Re: An Update on Our Outage
#30It’s great to hear that it wasn’t the result of something malicious (e.g. a hack). But this has to be one of the longest outages by a company this big ($50B market cap), at least in the past decade it seems?
What I love most then is that its the largest and it doesn't matter! In this market they will become an even larger company because competency has no correlation to how much money subscribers/advertisers will pay, and none of that has any correlation to what investors will pay! Anybody optimizing for competence is exchanging time for food and shelter and just gets to be reminded of the much harder game they are playi…