Live data from Hacker News

An Update on Our Outage

blog.roblox.com

181–190 of 235 posts

Re: An Update on Our Outage

#182
post #33

My heart goes out to the people who likely had to work crazy hours to fix this, but it really is wild that it was down for so long. What was the last service of this size that was down for 4 days? That is a failure in architecture that goes way beyond whatever the specific cause was here. That post-mortem is going to be a doozy.

Not necessarily directed at roblox but honestly I’m surprised this doesn’t happen more often given how many teams I see run software they don't understand or don’t even have access to its source code. Edit: but yeah must’ve been tough 72+hrs i hope their version of reliability team can use this to bash the support they need out of the management and not get scapegoated instead

So true. And also at the systems level. In most places I've seen, the #1 priority is hitting arbitrary executive feature/date goals, not maintaining robust systems. At some point, the shit will hit the fan, causing a "Why didn't you do perfectly the thing that wasn't a real priority?!?" reaction and a temporary lurch toward robustness. Although often the lurch will be less about actual robustness and more toward performative addition of control mechanisms like more layers of review.

I remember one large company I consulted for in the mid-aughts. Before I got there, they had an outage so severe it was on CNN and caused a short but notable dip in their stock price. By the time I got there, the ops people were absolutely dominant. All changes had to go through their review board, and woe be unto any project that they raised an eyebrow at.

Of course, the real problem was that developers were scheduled out 18 months in advance to work on a list of projects they hadn't been involved in estimating and where often they hadn't seen the code before being swapped onto the project. Dates were absolutely not allowed to slip, so it came out in frantic overwork, which meant a lot of bad/confusing code. Meaning more likelihood of failure and no time to think about systemic issues and impacts.

Re: An Update on Our Outage

#183
post #159

Earlier quoted context omitted.

Are you implying Netflix is not a tech company?

I definitely see Netflix as a Media Company, much like Disney. Disney and Pixar have great tech too, but they are not tech company in any shape or form. How Netflix got lumped into "FAANG" aka Big Tech is still a mystery to me.

NFLX was founded in 1997. They didn't have original content until 2013.

Before their original content I seriously considered shorting their stock. Their business model without original content was structurally flawed. They effectively had a maximum profit they could earn. Literally anything they did to increase their profit would be taken by the content producers they were forced to buy the licenses from.

Creating original content was a necessary step to survive just like pivoting from their original DVD mailing business to streaming. They're still a technology company though.

Re: An Update on Our Outage

#184
post #163
post #159

Earlier quoted context omitted.

I definitely see Netflix as a Media Company, much like Disney. Disney and Pixar have great tech too, but they are not tech company in any shape or form. How Netflix got lumped into "FAANG" aka Big Tech is still a mystery to me.

By that logic Facebook certainly wouldn't be a tech company either.

Facebook dont consider themselves as tech companies. But there aren't any other category that are clearly defined in which they fit in. Social Media isn't one.

Re: An Update on Our Outage

#185

Earlier quoted context omitted.

Not necessarily directed at roblox but honestly I’m surprised this doesn’t happen more often given how many teams I see run software they don't understand or don’t even have access to its source code. Edit: but yeah must’ve been tough 72+hrs i hope their version of reliability team can use this to bash the support they need out of the management and not get scapegoated instead

So true. And also at the systems level. In most places I've seen, the #1 priority is hitting arbitrary executive feature/date goals, not maintaining robust systems. At some point, the shit will hit the fan, causing a "Why didn't you do perfectly the thing that wasn't a real priority?!?" reaction and a temporary lurch toward robustness. Although often the lurch will be less about actual robustness and more toward perf…

I worked at a big fintech company about to go public in a SPAC worth $4 billion. They were a unicorn when I was there. Literally everything was in a GCP MySQL Database......and they refused to pay for a hot spare. They had backups but nothing for redundancy. We had downtime almost every week.

Re: An Update on Our Outage

#186
post #101

Earlier quoted context omitted.

Choice quotes from their PR piece: https://www.hashicorp.com/case-studies/roblox > We didn’t want to choose any technology that requires the company to drive deep expertise, almost to the point where you have to be a code contributor back into the project to get what you want. Nomad is just very easy to adopt. Better be damn sure you have your 24/7 vendor support contracts in order if and when shit does hit the fan.

> > We didn’t want to choose any technology that requires the company to drive deep expertise, That's a beautifully concise quote which neatly summarizes what contemporary IT values.

Abstracting things away is literally the entire point of programming.

Re: An Update on Our Outage

#187
post #33

My heart goes out to the people who likely had to work crazy hours to fix this, but it really is wild that it was down for so long. What was the last service of this size that was down for 4 days? That is a failure in architecture that goes way beyond whatever the specific cause was here. That post-mortem is going to be a doozy.

The immediate technical root cause and the true root cause (often down to poor decision making among the senior technical staff and management) are two different things, and I suspect the public and internal post-mortems both will focus on the former rather than the latter, sadly.

Re: An Update on Our Outage

#188
post #138

Earlier quoted context omitted.

I wonder if Engineers from Roblox are now worth a little more just because they have this experience. I also think it is time that people should take a look again at Chaos Engineering [1] from Netflix. It is sometimes ironic that the best technology often comes from companies that aren't a technology company at all. [1] https://principlesofchaos.org

Are you implying Netflix is not a tech company?

Since the idea of a "tech company" comes up so often, I'll just leave this here:

https://news.ycombinator.com/item?id=27693634

Re: An Update on Our Outage

#189

Earlier quoted context omitted.

> > We didn’t want to choose any technology that requires the company to drive deep expertise, That's a beautifully concise quote which neatly summarizes what contemporary IT values.

Abstracting things away is literally the entire point of programming.

Not understanding what the abstraction is for or what it is abstracting can be problematic when the abstraction breaks.

Wrong abstractions and too much abstraction can definitely both also be very bad.

Re: An Update on Our Outage

#190

Earlier quoted context omitted.

At my company there was a service outage that lasted 2-3 weeks for a specific feature we have. This was caused by everyone quitting and no one having any experience with this service. The rest of the application remained working so it wasn't so noticable to the outside world. But internally and for customers it was massive since it was the billing system that went down. There was another incident that took down every…

It’s been known for a while that human communication is the real impediment to technical development. Companies I’ve worked at that sucked had awful internal communication. It was all very friendly, but it was all platitudes and euphemisms, dumpster fire technology implementation. Phone and web apps are basically librarian work these days. If a businesses tech stack is having issues it’s human communication that’s th…

> Companies I’ve worked at that sucked had awful internal communication. It was all very friendly, but it was all platitudes and euphemisms, dumpster fire technology implementation.

Honestly, I think part of the problem is people going around talking so much about soft skills and how people don't like to be treated harshly that they've forgotten you need hard skills and nearly every profession where they need real leadership you get told where you screwed up and how to fix it. Literally, at the company I work for, people screw up and all you hear about is positive things

Post reply on HN