Live data from Hacker News

GCP Incidents

blog.railway.app

101–110 of 167 posts

Re: GCP Incidents

#101
post #29

Earlier quoted context omitted.

My guess is that whatever clever network optimizations that Google has are probably interfering with their traffic. By building their own network stack, they are skipping them and also wireguard might be better equipped to dealt with occasional faults as it built on udp which is inherently unreliable.

I have a direct cross connection to Google in a colocation facility (aka Dedicated Interconnect). One issue I found is that Google would randomly shuffle around their BGP routers which would cause BGP to flap and briefly losing all connectivity. When I raised this issue with support their answer was that this is expected behavior and we need to purchase a redundant connection. Mind you this isn’t cheap, we’re talking…

It’s about scale. Google was built for an order of magnitude greater scale where such reliability of a single link would be cost prohibitive.

However in general if you need high uptime, you’ll need multiple peering links. AWS and azure also recommend the same.

Re: GCP Incidents

#102
post #68

Earlier quoted context omitted.

>Azure Or you have big enterprise customers that have a grudge against Amazon and Google and refuse to use anything else.

Azure is great if you stick to the three golden oldies: DTU-based SQL, app service and service bus. Maybe table storage if you're feeling lucky. Anything else leads to pain and $$$, because it's likely that no one at Microsoft is being forced to use it.

vCore SQL is solid and predictable. Azure's vm offering is also highly reliable, and they host more linux workloads than Windows now.

Re: GCP Incidents

#103
post #64
post #62

I know AWS isn't cool or sexy, but shit works.

It’s sad to see people rediscovering that GCP is not a serious product over and over again.

That’s what I was thinking. Adopting GCP after years on AWS was eye-opening: I’d previously had a good impression but it was a constant cycle of “oh, we don’t have that - build your own” and issues which have been open for years full of customers asking and GCP PMs stalling.

Then the price increases started, making it even harder to defend paying more for less.

Re: GCP Incidents

#104
All these cloud service providers have bugs and issues.

But the problem with Google is that their support seems somehow disconnected from the real world. There is support, and they do respond to chats, calls, or emails. However, it often feels like I'm talking to someone who doesn't genuinely care about my concerns or do understand what I’m talking about.

Good support is hard to come by and hard to implement. So I really don't know what is missing in Google's support that exists in AWS support. Maybe because AWS support staff are trained to first put themselves in the customer's shoes and understand the problem from my perspective.

Re: GCP Incidents

#105
post #63

Earlier quoted context omitted.

Unnanounced changes too, there was a Firefox outage in 2022 due to GCP: https://hacks.mozilla.org/2022/02/retrospective-and-technica...

sorry but the blame here was 100% on Mozilla. No matter which http version, headers should always be treated as case-insensitive. Blanking anything on google here is just stupid. The problem was nih-syndrome and ignored the http spec.

Mozilla are entirely clear that this was their bug.

However, GCP changing the default under their infrastructure without prior warning was still unacceptable.

Operations work should (IMO must) be conducted with the expectation that any major change like that will expose existing bugs in deployed code.

(I've done enough ops work in my life that I'd love to say 'will potentially expose' but in practice there's always -something- that breaks and if I don't find it in the first 24h after a major change I'm going to spend the next two weeks waiting for the shoe drop to happen)

Re: GCP Incidents

#106
post #6

As someone who's into virtual worlds, and a user of Second Life, it's impressive to see how well those systems stay up. There hasn't been a total outage of Second Life in 5-10 years. Once Amazon's networking went down in a way that prevented new logins for a whole day, but existing logins remained. The 3D world, which has a lot of stuff going on even with no users around, continued to work. This is an extremely compl…

conways law affects the most those production support systems that everybody needs but nobody wants to pay for. Any boss will put their best talent on the critical revenue driving systems (e.g. adtech which is bounded by backend scalability). UI hasn’t really found a modern economic model yet (post Windows era), there are not research orgs dedicated to UI like there are infrastructure, databases, OS, ML etc

Re: GCP Incidents

#107
post #41

> In our experience, Google isn’t the place for reliable cloud compute, and it’s sure as heck not the place for reliable customer support. Always was, always will be. For them customers are always the last

GCP is the ugly stepchild of Google. They prioritize their own infra which unlike Aws doesn’t even run on gcp! It’s a joke. They don’t dogfood anything. All Google infra runs on separate systems (both) or dedicated deployment like their own spanner clusters. Google employees look down on gcp employees like second class citizens

> They prioritize their own infra which unlike Aws doesn’t even run on gcp!

I don't believe that to be the case, last I heard a year or two back the vast majority of it -does- run on a google GCP tenant account and what didn't was largely at least in the process of migration planning.

(my source here is "pillow talk with a senior GCP engineer" and I don't believe she had any reason to lie to me)

Re: GCP Incidents

#108

"reasons why Oxide has a business #12390"

An oxide rack has a minimum cost of something like 600k not including all the infra you need to run a rack, maintenance, and then needing to upgrade

Railway's bill was into the multiple millions per year at the very least so that doesn't necessarily rule it out.

Re: GCP Incidents

#109
post #6

As someone who's into virtual worlds, and a user of Second Life, it's impressive to see how well those systems stay up. There hasn't been a total outage of Second Life in 5-10 years. Once Amazon's networking went down in a way that prevented new logins for a whole day, but existing logins remained. The 3D world, which has a lot of stuff going on even with no users around, continued to work. This is an extremely compl…

[deleted]

Re: GCP Incidents

#110
post #104

All these cloud service providers have bugs and issues. But the problem with Google is that their support seems somehow disconnected from the real world. There is support, and they do respond to chats, calls, or emails. However, it often feels like I'm talking to someone who doesn't genuinely care about my concerns or do understand what I’m talking about. Good support is hard to come by and hard to implement. So I re…

[deleted]
Post reply on HN