Live data from Hacker News

Google Cloud networking issues in us-east1

status.cloud.google.com

181–190 of 341 posts

Re: Google Cloud networking issues in us-east1

#181
post #111

Earlier quoted context omitted.

Well, sure, if you hate your devops team and you want to make sure they can’t use any of the proprietary functionality of either provider. At which point, if you want to be managing a fleet of vanilla Linux boxes yourself, why use a cloud provider at all?

Why would you want to lock into a cloud provider? You're losing a lot of operational flexibility for less devops and sysaadmin work. You are really limiting your tech stack by using standardized things like Jenkins, Docker, K8, mqtt, kafka.

For the same reason you want "to lock in" (meaning use) any solution. You do not want to build or operate it yourself. Why don't you take this further? Why to use a water utility if you can just drill your own wells? Most businesses are better of on cloud because their core business is not to build and operate datacenters but provide services to their customers (on the top of datacenters running their apps).

Re: Google Cloud networking issues in us-east1

#182
post #3

Why so many problems at Google lately? Calendar down two weeks ago[0], and Google Cloud had a larger outage a month ago[1] [0]: https://news.ycombinator.com/item?id=20213092 [1]: https://news.ycombinator.com/item?id=20077421

> Why so many problems at Google lately?

I recently left Google to start a startup and now everything is falling apart.

Re: Google Cloud networking issues in us-east1

#183
post #160

Earlier quoted context omitted.

seriously, they've got a text field on the official status page, why not put the text boulos posted here in that instead of the meaningless text they've got there?

Can you expand on why you find it “meaningless”? As my other comment says, I’m not in SRE and the real people fixing it are trying their best to remediate the problem. I agree that the text I posted (with blessing from SRE!) gives you some more detail, but you can’t do anything differently with it, right? What about the new text do you prefer? (We’re happy to improve!)

I think the difference between your comment here and the info on the status page is that after reading your comment i feel like i know what's happening.

you're right, there's no additional actionable information there, the status page contains everything i actually need to know. but a bit more information makes me feel better. I guess the difference is your comment reassures me that you actually know what's going on. the status page text (prior to the 14:31 update) could equally mean "we've got this under control" or "shit's broken and we don't know why"

Re: Google Cloud networking issues in us-east1

#184
post #135

When choosing a big cloud provider people forget that it's many orders of magnitude more complicated to run something at Google scale then to maintain one single server. For example the whole Stack overflow website runs on one or two servers. World of Warcraft also used to run on one single (blade) server. Chances are one server will be good enough for most use cases. And if you don't want to have it in your closet t…

How can Stack Overflow run on a single server? Do you mean single cluster?

Re: Google Cloud networking issues in us-east1

#185

Earlier quoted context omitted.

It’s surprisingly hard to avoid shared fate links and it’s one of the things I would have thought google would be expert at.

It's not that hard. In India because of so much construction related digging cuts OFCs, we do the path planning quite well and our redundancies get tested quite regularly whether you want to or not.

> whether you want to or not

Accidents happen. Regularly. :D

Re: Google Cloud networking issues in us-east1

#186
post #160

Earlier quoted context omitted.

seriously, they've got a text field on the official status page, why not put the text boulos posted here in that instead of the meaningless text they've got there?

Can you expand on why you find it “meaningless”? As my other comment says, I’m not in SRE and the real people fixing it are trying their best to remediate the problem. I agree that the text I posted (with blessing from SRE!) gives you some more detail, but you can’t do anything differently with it, right? What about the new text do you prefer? (We’re happy to improve!)

"Can we meet up on Friday?"

"No" vs "No, I already have plans with X"

First case gives you all the information needed (denial), however in the second case I understand the situation much better. I wouldn't call the text on the status page meaningless though - it's pretty nice and concise already (which is what you want in a "crisis"). Just some brief description of the problem would be good, even though technically unnecessary.

Re: Google Cloud networking issues in us-east1

#187

Earlier quoted context omitted.

* You should not be locking yourself into proprietary functionality of a cloud provider unless you are deeply interested in what happened to Oracle customers getting raked over the coals happening to you. * DevOps teams can be multi-cloud relatively easy when using infrastructure as code tooling (Terraform, Packer, etc) and traditional DevOps practices * Why manage a fleet of vanilla boxes when you can use vanilla bo…

Proprietary managed services can save a lot of dev/setup/SRE time though. Many businesses have more pressing things to work on than spending dev time to prevent vendor lock-in.

Everyone spends their runway differently. Once you’re off the ground, derisk.

Re: Google Cloud networking issues in us-east1

#188
post #175

Earlier quoted context omitted.

It's not that hard. In India because of so much construction related digging cuts OFCs, we do the path planning quite well and our redundancies get tested quite regularly whether you want to or not.

It can be hard. Getting redundant separated paths under/over railroad tracks, for example, might require political power that not everyone has. Google, of course, has plenty.

> Getting redundant separated paths under/over railroad tracks, for example, might require political power that not everyone has. Google, of course, has plenty.

But Google's vendors might have less. One would hope that Google is auditing claims of independence from vendors at least somewhat, but at some level they have to rely on vendor representation and SLAs if they aren't going to do it all themselves.

Re: Google Cloud networking issues in us-east1

#189
> The disruptions with Google Cloud Networking and Load Balancing have been root caused to physical damage to multiple concurrent fiber bundles serving network paths in us-east1.

I am assuming some sort of construction zone at or nearby the facility and the backhoe operator dug in and accidently cut the cables?

Re: Google Cloud networking issues in us-east1

#190
post #88
post #12

Bad config push again?

Running a betting pool on cloud service outage root causes would be fairly fun. I'm going to guess load balancer cascading failures.

Nope. Physical destruction of fiber-optic cables is to blame, according to the GC status page. https://status.cloud.google.com/incident/cloud-networking/19...
Post reply on HN