Live data from Hacker News

Ongoing Incident in Google Cloud

status.cloud.google.com

71–80 of 115 posts

Re: Ongoing Incident in Google Cloud

#71
post #8

This demonstrates yet again why global configurations, global services, and global anycast VIP routing should be considered an anti pattern. gcp should be designed in a way where the term “global outage” isn’t a word in their vocabulary.

> This demonstrates yet again why global configurations, global services, and global anycast VIP routing should be considered an anti pattern.

And why enterprises clamoring for AWS to feature match Google's global stuff (theoretically making I.T. easier) instead of remaining regionally isolated (making I.T. actually more resilient, without extra work if I.T. operators can figure out infra-as-code patterns) should STFU and learn themselves some Terraform, Pulumi, or etc.

Also, AWS, if you're in this thread, stop with the recent cross-region coupling features already. Google's doing it wrong, explain that, and be patient, the market share will come back to you when they run out of the GCP subsidy dollars.

Re: Ongoing Incident in Google Cloud

#73
post #20

Earlier quoted context omitted.

Most of it is cellular or regional, but there are a few critical global services. The global network load balancing, network qos, and ddos prevention are more functional because they are global (i.e. you couldn't replace them with equivalent regional versions), but are often causes of issues like this. There was a push a few years ago to ensure global services had at least 99.999% uptime or make them regional. This w…

The pattern for past large google outages has been: 1. Some networking-related service has global, non-standard (compared to the rest of the company) configuration 2. The relevant VP is aware and has decided not to change it because that change is quoted as impossible 3. Some change elsewhere happens that assumes standard configuration 4. The networking service breaks and causes a global outage 5. VP is told to fix i…

Often "impossible" is based on constraints like "0 downtime" "100% planned rollout, rollback scenarios" etc.

These constraints get thrown to the wind when the downtime is already happening.

Re: Ongoing Incident in Google Cloud

#74
post #8

This demonstrates yet again why global configurations, global services, and global anycast VIP routing should be considered an anti pattern. gcp should be designed in a way where the term “global outage” isn’t a word in their vocabulary.

> This demonstrates yet again why global configurations, global services, and global anycast VIP routing should be considered an anti pattern. And why enterprises clamoring for AWS to feature match Google's global stuff (theoretically making I.T. easier) instead of remaining regionally isolated (making I.T. actually more resilient, without extra work if I.T. operators can figure out infra-as-code patterns) should STF…

You really want to go through every region to find what VMs are running? Why can this not be a single page with all VMs listed?

Re: Ongoing Incident in Google Cloud

#76
As has happened many times throughout history (back to mainframes and thin clients of the 90s) there are swings/trends in how infrastructure is hosted.

Listening to the “All In Podcast” yesterday even those guys were talking about revenue drops in the big cloud services and noting we’re currently in the midst of a swing back to self-hosting/co-location/whatever thinking and migrations out.

IMHO those building greenfield solution today should take a hard look at whether the default approach from the last ~10 years “of course you build in $BIGCLOUD” makes sense for the application - in many cases it does not.

It also has the added benefit of de-centralizing the internet a bit (even if only a little).

Re: Ongoing Incident in Google Cloud

#77

As has happened many times throughout history (back to mainframes and thin clients of the 90s) there are swings/trends in how infrastructure is hosted. Listening to the “All In Podcast” yesterday even those guys were talking about revenue drops in the big cloud services and noting we’re currently in the midst of a swing back to self-hosting/co-location/whatever thinking and migrations out. IMHO those building greenfi…

Revenue drop? Google Cloud is still growing 30-40% year on year.

Re: Ongoing Incident in Google Cloud

#78
post #77

As has happened many times throughout history (back to mainframes and thin clients of the 90s) there are swings/trends in how infrastructure is hosted. Listening to the “All In Podcast” yesterday even those guys were talking about revenue drops in the big cloud services and noting we’re currently in the midst of a swing back to self-hosting/co-location/whatever thinking and migrations out. IMHO those building greenfi…

Revenue drop? Google Cloud is still growing 30-40% year on year.

AWS also had 20% revenue growth last quarter.

Re: Ongoing Incident in Google Cloud

#79

As has happened many times throughout history (back to mainframes and thin clients of the 90s) there are swings/trends in how infrastructure is hosted. Listening to the “All In Podcast” yesterday even those guys were talking about revenue drops in the big cloud services and noting we’re currently in the midst of a swing back to self-hosting/co-location/whatever thinking and migrations out. IMHO those building greenfi…

[deleted]

Re: Ongoing Incident in Google Cloud

#80

As has happened many times throughout history (back to mainframes and thin clients of the 90s) there are swings/trends in how infrastructure is hosted. Listening to the “All In Podcast” yesterday even those guys were talking about revenue drops in the big cloud services and noting we’re currently in the midst of a swing back to self-hosting/co-location/whatever thinking and migrations out. IMHO those building greenfi…

All in podcast mentioned growth slowing, but not revenue dropping.
Post reply on HN