This demonstrates yet again why global configurations, global services, and global anycast VIP routing should be considered an anti pattern. gcp should be designed in a way where the term “global outage” isn’t a word in their vocabulary.
Ongoing Incident in Google Cloud
21–30 of 115 posts
Re: Ongoing Incident in Google Cloud
#22<3 To the engineers trying to fix it at the moment.
Re: Ongoing Incident in Google Cloud
#23https://packages.cloud.google.com/apt/doc/apt-key.gpg Even the public apt key for signing Google's cloud packages is unavailable (returns 500 for me). This is insane
Re: Ongoing Incident in Google Cloud
#24Earlier quoted context omitted.
My knowledge level: can use AWS console to do How much more work would Google create for themselves if they had not globalized their stack? Are we talking something like 5 subsets to manage instead of 1?
Most of it is cellular or regional, but there are a few critical global services. The global network load balancing, network qos, and ddos prevention are more functional because they are global (i.e. you couldn't replace them with equivalent regional versions), but are often causes of issues like this. There was a push a few years ago to ensure global services had at least 99.999% uptime or make them regional. This w…
1. Some networking-related service has global, non-standard (compared to the rest of the company) configuration
2. The relevant VP is aware and has decided not to change it because that change is quoted as impossible
3. Some change elsewhere happens that assumes standard configuration
4. The networking service breaks and causes a global outage
5. VP is told to fix it
6. Fix rolls out in weeks, because it wasn't as hard as they said before
Re: Ongoing Incident in Google Cloud
#25Re: Ongoing Incident in Google Cloud
#26This demonstrates yet again why global configurations, global services, and global anycast VIP routing should be considered an anti pattern. gcp should be designed in a way where the term “global outage” isn’t a word in their vocabulary.
Re: Ongoing Incident in Google Cloud
#27Re: Ongoing Incident in Google Cloud
#28Not great, not terrible.
Re: Ongoing Incident in Google Cloud
#29Outages at the hyperscalers can have a huge blast radius, is anyone encountering other services with outages because they're built on GCP?
Re: Ongoing Incident in Google Cloud
#30This demonstrates yet again why global configurations, global services, and global anycast VIP routing should be considered an anti pattern. gcp should be designed in a way where the term “global outage” isn’t a word in their vocabulary.
Of course, at Google scale 'partial' is still very big.