Live data from Hacker News

Meta outage

metastatus.com

721–730 of 902 posts

Re: Meta outage

#721

Earlier quoted context omitted.

Is it a false positive, though? The data shows there was an outage. We would need more evidence to conclude hundreds of users, at that 1 spike, weren't actually having issues. In other words, we have hundreds of people saying there was an outage, and 1 person saying there wasn't. That's a problem AWS needs to resolve, regardless of what they think might be the root cause. If the users weren't experiencing any issues…

> Is it a false positive, though? Yes. AWS was not down this morning. > In other words, we have hundreds of people saying there was an outage, and 1 person saying there wasn't. We have hundreds of millions using AWS and AWS-backed services successfully this morning. I'm out.

Hundreds of users, representing more users who didn't bother reporting, say they experienced issues when interacting with AWS this morning, so we'll need better evidence to the contrary to conclude otherwise.

The fact that some people accessed AWS without reporting issues does not mean that all people did. For those who had issues, AWS is responsible for dealing with those perceptions.

Indeed, it could have been a fault that affected a subset of users, for example 1 service in 1 availability zone. That's still an outage in the eyes of users, which AWS is responsible for managing. It could have been an issue with a route from 1 ISP. That's still an outage in the eyes of users, which AWS is responsible for managing.

An even better example is the DownDetector page for Facebook, with hundreds of thousands of reports. Do we really think there's no correlation between what DownDetector reports and what users experience?

tl;dr: what users think about your site is more important than both what you think about your site and the reality of your site, and you should be tracking it.

Re: Meta outage

#722

Earlier quoted context omitted.

That's the natural endgame of the "user-facing services must not stop, if something they depend upon stops, they must only degrade" philosophy.

I heard for a while Netflix would fail open if auth was unavailable. Like it’s just movies just let em see it. Facebook data is more sensitive. Not so much the data people go there to see, cool memes that their friends liked, but the list of friends and interests. Other places I worked had the ability for Ops to push out a change saying the site was down for maintenance. After a while we stopped using it and just too…

Failing open is maybe ok. Telling everybody on the world their account doesn't exist anymore isn't.

Re: Meta outage

#723
post #292

I was considering buying the Quest 3 this morning, this outage is timely. The fact that I can't use my perfectly working headset to play an offline game because facebook is down, makes me wonder if I should go for a different provider. Any recommendations? Excluding apple vision pro since it is too expensive.

The Sony PSVR2 is getting PC support by end of year apparently.

Re: Meta outage

#724
Since my dedicated post about this got no response, I want to use this opportunity to ask:

WHERE the HELL does Facebook store tracking data on my iPhone?

It shows my previous account even after I delete the app, clear the cache and KeyChain, disable iCloud Drive, AND sign out of iCloud??

Why can't I see where this data is stored? Same for TikTok.

WHY does Apple, parading around as a pompous paragon of privacy, allow this bullshit?

Re: Meta outage

#726
post #292

I was considering buying the Quest 3 this morning, this outage is timely. The fact that I can't use my perfectly working headset to play an offline game because facebook is down, makes me wonder if I should go for a different provider. Any recommendations? Excluding apple vision pro since it is too expensive.

The Sony PSVR2 is getting PC support by end of year apparently.

This is pretty cool. I might get one if I’m convinced it’s better than my quest3s display link.

Re: Meta outage

#727

Earlier quoted context omitted.

[flagged]

I can guarantee you with 100% confidence from experience that the call centers for AT&T, T-Mobile, Comcast, etc. are all blowing up right now because of users who assume that if the Instagram app isn’t loading it means the “wifi” is broken. Also keep in mind “wifi” doesn’t mean 802.11, it means “anything related to the internet” up to and including 4g/5g and Ethernet.

Ok great. How does that equate to having down detector filter reports?

Re: Meta outage

#728
post #438

Earlier quoted context omitted.

At the scale of Meta, "down" is a nuanced concept. You are very unlikely to get every piece of functionality seizing up at once. What you are likely to get is some services ceasing to function and other services doing error-handling. For example, if the service that authenticates a user stops working but the service that shows the login form works, then you get a complex interaction. The resulting messaging - and thu…

Isn't this just the standard problem of reporting useful error messages? Like, yes, there are academic situations where you can't distinguish between two possible error sources, but the vast majority of insufficiently informative error messages in the real world arise because low effort was applied to doing so.

Yes and no.

Yes, with the additions of sheer scale, a vast number of services, multiple layers, and the difficulty of defining "down" added in. I think the difficulty of reporting useful error messages is proportional to the number of places an error can reasonably happen and the number of connections it can happen over, and by any metric Meta's got a lot of those.

No, in that detecting when you should be reporting a useful error message is itself a complex problem. If a service you call gives you a nonsense response, what do you surface to the user? If a service times out, what do you report? How do you do all this without confusing, intimidating, and terrifying users to whom the phrase "service timeout" is technobabble?

Re: Meta outage

#729
post #391

Looking at the Downdetector home page [1], it looks like many more services are having outages, not just the ones owned by Meta, including: - Google - YouTube - Google Play - T-Mobile - X (Twitter) - Discord - TikTok - Pokemon Go - Snapchat It looks like they all have the same failure point. [1] https://downdetector.com

There is no consistent scale on that graph, so any local maxima of reports received would look similar to any other.

I made that same mistake after seeing someone post an unlabeled set of graphs to a Slack. The Google peak reported outages is about 0.25% of the Facebook peak. It seems reasonable some people just made a mistake.

Re: Meta outage

#730

Earlier quoted context omitted.

[flagged]

> do you really think there are masses of people who can’t tell the difference between a single sign on service being down and individual sites being down and reporting it to downdetector? Have you never seen The Website Is Down? https://www.youtube.com/watch?v=uRGljemfwUE The answer is: way more people than a software developer might think. Ask anyone in IT, or go to anywhere bugs are reported and read a handful.

yeah that’s my point. No one who arranges icons by penis is taking time to go file a report on down detector.
Post reply on HN