Live data from Hacker News

Ongoing Incident in Google Cloud

status.cloud.google.com

91–100 of 115 posts

Re: Ongoing Incident in Google Cloud

#91

Earlier quoted context omitted.

As others have mentioned, there was no revenue drop, there's been a reduction in growth. AWS's 20% growth rate is still very respectable, more than double the 9% growth rate the company had overall. I would be hesitant to attribute slowed growth to a return to self hosting, it's much more likely that it's caused by companies dialing back their cloud growth after spending a few years going ham digitizing everything du…

parent still has a very strong point considering that a drop in growth (not revenue) quickly translates in projects / features being cancelled. That's a good thing to FailFast from a start-up pov but when me as a start-up needs to make a bet about building on top of certain features this adds to my cost/benefit calculation when deciding if I want to jump on new features (device-shadows, digital-twins, or whatever els…

But there are so many degrees of ratcheting back cloud costs before we get back to self-hosted.

Sure, companies are probably less interested in wacky new cloud features then they were before, but that means going back to basics like EC2 and RDS, which do function like utilities, not going back to their own data centers.

Re: Ongoing Incident in Google Cloud

#92
post #85

Earlier quoted context omitted.

AWS is a 75 bln a year business still growing 20%+ YoY. It’ll break 100 bln this year. I would examine the numbers yourself.

I have - and the numbers show that much of the big cloud growth is in AI services. The "we need to throw in AI somewhere" concurrent trend is heavily bolstering what would other wise be much more drastic retractions in growth. I would argue as the AI trend (eventually) wanes and many AI startups and projects within existing companies inevitably eventually fail to materialize the much longer and more general trend of…

> I have - and the numbers show that much of the big cloud growth is in AI services. The "we need to throw in AI somewhere" concurrent trend is heavily bolstering what would other wise be much more drastic retractions in growth.

Can you share where you got this? Which numbers? I didn't think AWS (or any cloud provider) released details of their operation at that level of granularity.

Re: Ongoing Incident in Google Cloud

#93

Earlier quoted context omitted.

> This demonstrates yet again why global configurations, global services, and global anycast VIP routing should be considered an anti pattern. And why enterprises clamoring for AWS to feature match Google's global stuff (theoretically making I.T. easier) instead of remaining regionally isolated (making I.T. actually more resilient, without extra work if I.T. operators can figure out infra-as-code patterns) should STF…

You really want to go through every region to find what VMs are running? Why can this not be a single page with all VMs listed?

The AWS console already has a single page where you can see how many EC2/networking resources you have in every region.

Re: Ongoing Incident in Google Cloud

#94

Earlier quoted context omitted.

As others have mentioned, there was no revenue drop, there's been a reduction in growth. AWS's 20% growth rate is still very respectable, more than double the 9% growth rate the company had overall. I would be hesitant to attribute slowed growth to a return to self hosting, it's much more likely that it's caused by companies dialing back their cloud growth after spending a few years going ham digitizing everything du…

Ah yes, sorry, slower than expected growth was the data point. In my defense I had a screaming toddler in the car! That said I think the point generally remains - one could argue slower than expected growth in cloud services is a revenue drop (in a way) vs expectations. The market responded accordingly[0] - "However, Azure growth is decelerating." Note that this is all including the explosion in "2023 AI hotness" whi…

I think you're still jumping to conclusions to think that the ratcheting back is going to take any significant portion of the market all the way back to self-hosted. I suspect that companies are less willing to invest in fancy new platform features that drive more revenue than VPSs and managed DBs, but I have a very hard time believing that EC2 or RDS are flagging.

Re: Ongoing Incident in Google Cloud

#95
post #8

This demonstrates yet again why global configurations, global services, and global anycast VIP routing should be considered an anti pattern. gcp should be designed in a way where the term “global outage” isn’t a word in their vocabulary.

  > gcp should be designed in a way where the term “global outage” isn’t a word in their vocabulary.
If that's what you really need, then distribute your assets across GCP, AWS, and DO. That likely means not using any cloud-specific features such as Lambda. AWS is actually really good in this regard, as SES and RDS are easily copied to regular instances in other cloud providers, that possibly wrap some cloud-specific feature themselves.

Re: Ongoing Incident in Google Cloud

#96

Earlier quoted context omitted.

As others have mentioned, there was no revenue drop, there's been a reduction in growth. AWS's 20% growth rate is still very respectable, more than double the 9% growth rate the company had overall. I would be hesitant to attribute slowed growth to a return to self hosting, it's much more likely that it's caused by companies dialing back their cloud growth after spending a few years going ham digitizing everything du…

Ah yes, sorry, slower than expected growth was the data point. In my defense I had a screaming toddler in the car! That said I think the point generally remains - one could argue slower than expected growth in cloud services is a revenue drop (in a way) vs expectations. The market responded accordingly[0] - "However, Azure growth is decelerating." Note that this is all including the explosion in "2023 AI hotness" whi…

In isolated cases companies can reduce costs by self hosting. Usually this is a combination of very specialized requirements or shockingly technically competent early employees or founders.

However even in these exceptional cases there are hidden costs that will likely arise.

For most companies the very concept of self hosting is comical. This is a one way train.

Re: Ongoing Incident in Google Cloud

#97

Earlier quoted context omitted.

The pattern for past large google outages has been: 1. Some networking-related service has global, non-standard (compared to the rest of the company) configuration 2. The relevant VP is aware and has decided not to change it because that change is quoted as impossible 3. Some change elsewhere happens that assumes standard configuration 4. The networking service breaks and causes a global outage 5. VP is told to fix i…

Often "impossible" is based on constraints like "0 downtime" "100% planned rollout, rollback scenarios" etc. These constraints get thrown to the wind when the downtime is already happening.

I was being a bit hyperbolic, but this is the real reason. However, the VPs in question often have the authority to approve changes that don't have rollback scenarios (for example), they just don't until the shit hits the fan.

Re: Ongoing Incident in Google Cloud

#98

Earlier quoted context omitted.

They had many many global outages through the years so that’s evidently not true. GCLB, iam, gcs and probably more Im missing just of the top of my head. Then there’s constant stream of regional networking borks where your latency is suddenly 5x which are not “global” but affect multiple regions

Anecdata, but in my experience Google Cloud has been MUCH more solid than my time spent on AWS.

Historically that has not been my experience at all tho tbf gcp has cleaned up their act substantially in the past 1-2 years

Re: Ongoing Incident in Google Cloud

#99
post #21
post #8

This demonstrates yet again why global configurations, global services, and global anycast VIP routing should be considered an anti pattern. gcp should be designed in a way where the term “global outage” isn’t a word in their vocabulary.

You can't really have 30+ fully independent regions running their own stack with different versions of apps and separate secrets, IP/routing and certificates in each. At some point you have to unify or it becomes either unmanageable or inconsistent.

How does this fit in with upcoming EU data sovereignty laws?

Re: Ongoing Incident in Google Cloud

#100

As has happened many times throughout history (back to mainframes and thin clients of the 90s) there are swings/trends in how infrastructure is hosted. Listening to the “All In Podcast” yesterday even those guys were talking about revenue drops in the big cloud services and noting we’re currently in the midst of a swing back to self-hosting/co-location/whatever thinking and migrations out. IMHO those building greenfi…

> IMHO those building greenfield solution today should take a hard look at whether the default approach from the last ~10 years “of course you build in $BIGCLOUD” makes sense for the application - in many cases it does not. When one buys a house, they should take a hard loo at whether the default approach of paying for utilities makes sense, versus generating their own power. While that's a bit snarky, the reasoning…

> When one buys a house, they should take a hard loo at whether the default approach of paying for utilities makes sense, versus generating their own power.

Yes, people do. They install solar panels and use them to generate at least some of their own power. Near future battery tech might allow them to generate all of it if they get enough sunlight, in which case this will become a genuine question to answer: how much to install and maintain the panels and batteries over their lifetime, vs expected cost of purchasing power from utilities.

In a similar manner, cloud vs self hosting is a valid consideration that changes over time. We now have docker and similar tools which make managing your own infrastructure much easier than it was ten years ago. I fully expect even better tools will come out in the future so this consideration does change over time. Maybe in another ten years there'll be almost no benefit to using the cloud (except maybe as a CDN).

Post reply on HN