Live data from Hacker News

Ongoing Incident in Google Cloud

status.cloud.google.com

101–110 of 115 posts

Re: Ongoing Incident in Google Cloud

#101
post #19

Earlier quoted context omitted.

> gcp should be designed in a way where the term “global outage” isn’t a word in their vocabulary. As I understand it, GCP is already designed to make global outages impossible. Obviously this outage shows that they messed up somehow and some global point of failure still remains. Looking forward to the post-mortem.

They had many many global outages through the years so that’s evidently not true. GCLB, iam, gcs and probably more Im missing just of the top of my head. Then there’s constant stream of regional networking borks where your latency is suddenly 5x which are not “global” but affect multiple regions

They had lots of global outages in past years, but in recent years they have become increasingly rare, presumably because of a move away from global points of failure

Re: Ongoing Incident in Google Cloud

#102
post #67
post #21

Earlier quoted context omitted.

You can't really have 30+ fully independent regions running their own stack with different versions of apps and separate secrets, IP/routing and certificates in each. At some point you have to unify or it becomes either unmanageable or inconsistent.

But you can have 3. Why did you choose 30? In my company we are split in 3, US, EU, APAC, and we have the same issue with global outage for stuff we could have just managed regionally. For all the savings of the global architecture, they disappear each minute a client is down on a global outage because a guy thousands of kms away messed up. You dont have to unify, at all. You dont unify with your competitors, and the…

GCP needs to support 30 regions because... they're a cloud provider.

Re: Ongoing Incident in Google Cloud

#103
post #59

Earlier quoted context omitted.

https://aws.amazon.com/builders-library/automating-safe-hand...

AWS US east 1 had significant downtime last year so I'm not sure what you're trying to say with that link. Would you mind expanding on your thoughts?

One region failing (especially us-east-1) is common, but it's very rare to see an AWS global outage.

Re: Ongoing Incident in Google Cloud

#104
post #67

Earlier quoted context omitted.

But you can have 3. Why did you choose 30? In my company we are split in 3, US, EU, APAC, and we have the same issue with global outage for stuff we could have just managed regionally. For all the savings of the global architecture, they disappear each minute a client is down on a global outage because a guy thousands of kms away messed up. You dont have to unify, at all. You dont unify with your competitors, and the…

GCP needs to support 30 regions because... they're a cloud provider.

Then they can do a glocal model with 10 regions grouped into 3 semi global groups, so when there s a global outage, it can only be on one of these ?

Re: Ongoing Incident in Google Cloud

#105
post #59

Earlier quoted context omitted.

AWS US east 1 had significant downtime last year so I'm not sure what you're trying to say with that link. Would you mind expanding on your thoughts?

One region failing (especially us-east-1) is common, but it's very rare to see an AWS global outage.

This. us-east-1 is the oldest region IIRC and it has its share of issues. Back when I used to work mostly on AWS zonal outages used to happen once in a while, but entire regions were rare, forget global outages.

The global outage thing seems to be a consistent "feature" of GCP - how are we supposed to architect our deployments if the regional isolation model is not a bulwark against high availability on GCP?

Re: Ongoing Incident in Google Cloud

#106
post #8

This demonstrates yet again why global configurations, global services, and global anycast VIP routing should be considered an anti pattern. gcp should be designed in a way where the term “global outage” isn’t a word in their vocabulary.

> gcp should be designed in a way where the term “global outage” isn’t a word in their vocabulary. If that's what you really need, then distribute your assets across GCP, AWS, and DO. That likely means not using any cloud-specific features such as Lambda. AWS is actually really good in this regard, as SES and RDS are easily copied to regular instances in other cloud providers, that possibly wrap some cloud-specific f…

But....cost.

Re: Ongoing Incident in Google Cloud

#107
post #70

Earlier quoted context omitted.

Ah so the rotation rotates to match the current rotation. Very smart.

I like the mental image of this being a very precise matching -- as the sun traces across the sky, the responsibility of on-call passes from desk to desk, town to town, country to country; two engineers on a boat in the Atlantic race to keep up with their rotation...

The logical conclusion is SREpiercer

Re: Ongoing Incident in Google Cloud

#108

As has happened many times throughout history (back to mainframes and thin clients of the 90s) there are swings/trends in how infrastructure is hosted. Listening to the “All In Podcast” yesterday even those guys were talking about revenue drops in the big cloud services and noting we’re currently in the midst of a swing back to self-hosting/co-location/whatever thinking and migrations out. IMHO those building greenfi…

The parent comment is neither factual nor advisable.

You build greenfield in cloud precisely because it is greenfield and the utilization isn't well understood. Cloud options let you adjust and experiment quickly. Once a workload is well understood it's a good candidate for optimization, including a move to self managed hardware / on prem.

Buying hardware is a great option once you actually understand the utilization of your product. Just make sure you also have competent operators.

Re: Ongoing Incident in Google Cloud

#109
post #106

Earlier quoted context omitted.

> gcp should be designed in a way where the term “global outage” isn’t a word in their vocabulary. If that's what you really need, then distribute your assets across GCP, AWS, and DO. That likely means not using any cloud-specific features such as Lambda. AWS is actually really good in this regard, as SES and RDS are easily copied to regular instances in other cloud providers, that possibly wrap some cloud-specific f…

But....cost.

Reliable, cheap, or ....? Pick N-1.

Re: Ongoing Incident in Google Cloud

#110

Earlier quoted context omitted.

Ah yes, sorry, slower than expected growth was the data point. In my defense I had a screaming toddler in the car! That said I think the point generally remains - one could argue slower than expected growth in cloud services is a revenue drop (in a way) vs expectations. The market responded accordingly[0] - "However, Azure growth is decelerating." Note that this is all including the explosion in "2023 AI hotness" whi…

I think you're still jumping to conclusions to think that the ratcheting back is going to take any significant portion of the market all the way back to self-hosted. I suspect that companies are less willing to invest in fancy new platform features that drive more revenue than VPSs and managed DBs, but I have a very hard time believing that EC2 or RDS are flagging.

I should have been more clear on this - in terms of total install base I don't know that it's going to be "significant" in terms of customer count.

However, I do think it will be at least "noticeable" in terms of individual customers with large spend. Total GCP revenue in 2022 was roughly 65 billion and Snap leaving alone is 1.5% of total revenue.

Especially looking at ML cases where cloud GPU pricing is wildly expensive - retail on-demand instance A100 pricing is at least $3/hr which practically speaking with the AWS pricing model is can be twice that all-in. This is for an instance with 32GB of RAM and 8 VCPUs - which for a lot of A100 use cases is useless. Need 32 vCPU and 256 GB of RAM? That's more like $20/hr.

A single A100 machine that's above and beyond more capable can be had from Dell for roughly $50k, which even factoring in hosting based on colo pricing I've seen has an ROI of ~15 months for constant usage. For the equivalent hardware (and still vastly improved performance - 32vCPU and 256GB of RAM) that ROI gets to less than six months.

Yes, the A100 is typically used for training (and cloud definitely still makes sense there) but more and more models require the performance and VRAM of a V100/H100 for inference (24/365 availability). Do it at any kind of scale/redundancy and ROI catches up even faster. An equivalent to this approach is reserved pricing, which over the 1-3yr term of a lease vs. reserved instance self-hosting becomes almost comically more cost and performance effective. With the extra benefit of actually being more flexible.

Financing and leasing is readily available and with various tax incentives (like Section 179 leasing) you can pretty quickly pay for a FTE to manage the infra for you - which is probably a wash anyway because at any kind of "real" scale or complexity you almost certainly already have dedicated human resources just to manage cloud. You don't even ever need for an employee to go to the hosting facility because most will rack and provision your hardware for free. Combined with remote hands and standard warranty support any (in my experience very rare) hardware failures just get handled.

I should note that this model almost eliminates the tendency for cloud spend to balloon to many X anticipated/budgeted spend - the all too common story of "sticker shock" from clouds on bandwidth alone that cloud has ridiculous markups on. Colo pricing and leases are fixed cost (with all you can eat port speed bandwidth included or so cheap at 95th percentile billing it's practically a rounding error).

I have significant experience at CTO level with both approaches (and hybrid, of course). In many situations the benefits of "self-hosting" vs cloud are dramatic.

The extremely effective marketing that has created and perpetuated an industry wide fear of self-hosting and hardware (especially with the "always cloud always" generation) is fading. I think the uptime and reliability promises of cloud are also fading - this thread started off with discussion of yet-another cloud outage. My background is in healthcare and telecom and I'm shocked at the cavalier attitude of just accepting these outages and being down while standing around helpless wondering when your big cloud will acknowledge, communicate, and resolve them. A few machines in a few rack units of space across a couple of facilities generally trounces cloud in reliability and uptime.

I love HN and the overall knowledge and quality of discussion here but when it comes to hardware and self-hosting many have completely drunk the cloud Kool-Aid and have zero experience with self-hosting - so no idea what they're talking about. Not saying you personally, just generally.

Post reply on HN