Live data from Hacker News

An Update on Our Outage

blog.roblox.com

171–180 of 235 posts

Re: An Update on Our Outage

#171

Earlier quoted context omitted.

I think it happens all the time. There just aren't that many services, relatively speaking, that you'd hear about when this happens. The more-popular services (that you'd hear about) probably already have a bias toward having their stuff better in order, since it's often also really just a matter of making the right investments.

> really just a matter of making the right investments Right. But the person deciding on those investments is a product owner who's looking to release new features because that's what management wants. At least that's my experience more often than not. My guess is that the bias might be the opposite. Maybe these companies become more popular because they are reliable and have the mindset to focus on that.

By the time a company gets large enough to be a household name, they have usually been burned a few times and have someone technical who can push back against short-termism from POs.

Re: An Update on Our Outage

#172
post #138

Earlier quoted context omitted.

I wonder if Engineers from Roblox are now worth a little more just because they have this experience. I also think it is time that people should take a look again at Chaos Engineering [1] from Netflix. It is sometimes ironic that the best technology often comes from companies that aren't a technology company at all. [1] https://principlesofchaos.org

And I wonder if their tech management is worth less for this clusterf** to happen during their oversight?

This is how I felt reading his comment. Were Volkswagen emission engineers and the captain of the Exxon Valdez worth more after they screwed up?

Re: An Update on Our Outage

#173
post #26

As most of the Roblox community is aware, we recently experienced an extended outage across our platform. We are sorry for the length of time it took us to restore service. A key value at Roblox is “Respect the Community,” and in this case, we apologize for the inconvenience to our community. On Thursday afternoon, October 28th, users began having trouble connecting with our platform. This immediately became our high…

Thank you for posting the text of the blog post. However, I don't understand your TLDR. They clearly were able to identify the root cause, and that is what allowed them to restore service.

People use "root cause" to mean the true underlying driver. You can often restore service without knowing the true root cause. In this example, let's pretend true root cause is a memory leak that takes X days to crash Consul. It's possible they don't know that yet, and just shut down, cleaned up logs and temporary storage, added some capacity, and started back up in a region-by-region ramp up.

Re: An Update on Our Outage

#174

Earlier quoted context omitted.

I mean I'd gladly pull an all-nighter if I were a millionaire thanks to the thing I'm fixing. I think. Then I'd quit and do something leisurely.

Over what time frame would you tolerate poor working conditions to make a million bucks? Four year equity vesting? A million dollars in 2025 too, not a million dollars in 2010.

Just one million? About 6 months, as I currently earn over $2MM per year from my options alone at the company I work for thanks to their IPO. Once those dry up it's going to be a hard sell for me to stay. Maybe I'll coast for a year on cash + RSU salary (which is much less than salary plus pre-IPO options) to top off my $5MM investment portfolio and ensure enough of it is invested to produce an income, but after that? No thanks.

Re: An Update on Our Outage

#175
post #136

Earlier quoted context omitted.

This is a benefit of a global company, not working remote. You could do all that stuff in multiple offices around the world.

In a global but not-remote company the team responsible for the particular service that failed is probably still concentrated in one office. Bringing in people from other offices who aren't familiar with the problem service probably isn't that helpful.

You can have knowledge silos in either kind of company. We didn't magically get better at spreading work around or writing documentation just because we went remote.

Re: An Update on Our Outage

#176
post #82
post #33

My heart goes out to the people who likely had to work crazy hours to fix this, but it really is wild that it was down for so long. What was the last service of this size that was down for 4 days? That is a failure in architecture that goes way beyond whatever the specific cause was here. That post-mortem is going to be a doozy.

In the game industry, working yourself to death for 4 days, that's called a hot fix ;)

People working on the cloud infra are not the people working on games.

Re: An Update on Our Outage

#177

Earlier quoted context omitted.

Not necessarily directed at roblox but honestly I’m surprised this doesn’t happen more often given how many teams I see run software they don't understand or don’t even have access to its source code. Edit: but yeah must’ve been tough 72+hrs i hope their version of reliability team can use this to bash the support they need out of the management and not get scapegoated instead

At my company there was a service outage that lasted 2-3 weeks for a specific feature we have. This was caused by everyone quitting and no one having any experience with this service. The rest of the application remained working so it wasn't so noticable to the outside world. But internally and for customers it was massive since it was the billing system that went down. There was another incident that took down every…

It’s been known for a while that human communication is the real impediment to technical development.

Companies I’ve worked at that sucked had awful internal communication. It was all very friendly, but it was all platitudes and euphemisms, dumpster fire technology implementation.

Phone and web apps are basically librarian work these days. If a businesses tech stack is having issues it’s human communication that’s the real problem.

Re: An Update on Our Outage

#178

Earlier quoted context omitted.

I think it happens all the time. There just aren't that many services, relatively speaking, that you'd hear about when this happens. The more-popular services (that you'd hear about) probably already have a bias toward having their stuff better in order, since it's often also really just a matter of making the right investments.

> really just a matter of making the right investments Right. But the person deciding on those investments is a product owner who's looking to release new features because that's what management wants. At least that's my experience more often than not. My guess is that the bias might be the opposite. Maybe these companies become more popular because they are reliable and have the mindset to focus on that.

Either you listen to and trust your people or your don't.

Re: An Update on Our Outage

#179
post #33

My heart goes out to the people who likely had to work crazy hours to fix this, but it really is wild that it was down for so long. What was the last service of this size that was down for 4 days? That is a failure in architecture that goes way beyond whatever the specific cause was here. That post-mortem is going to be a doozy.

Not necessarily directed at roblox but honestly I’m surprised this doesn’t happen more often given how many teams I see run software they don't understand or don’t even have access to its source code. Edit: but yeah must’ve been tough 72+hrs i hope their version of reliability team can use this to bash the support they need out of the management and not get scapegoated instead

[deleted]

Re: An Update on Our Outage

#180
post #163

Earlier quoted context omitted.

By that logic Facebook certainly wouldn't be a tech company either.

Why not? Didn't Facebook create tools like GraphQL and React?

Netflix has created tools and open source libraries too.

RxJava was opensourced by Netflix, the defacto standard for reactive programming in Java. They also open sourced a lot of their work for working with avif files, which has managed to find its way into Chrome, among other projects.

Among other things.

Post reply on HN