AWS CloudWatch has an option to show the trend and what it will be like after x-period. Doesn't Azure have such options so that engineers can predict to scale better? Seems like engineers are not ready for this per postmortem
The August 17 outage
31–40 of 804 posts
Re: The August 17 outage
#32What I read from this is: * Scaling is hard, we don't have enough capacity * We give away a shitton of compute for free * I have to talk about Azure not being a steaming pile of poop, otherwise my bonus will get tweaked downward in the next comp cycle.
Notice there's nothing about paid customers, I'll add in what they are missing:
Paid customers: Go F*ck yourself, you don't pays us enough to be an interesting line item compared to windows server.
Re: The August 17 outage
#33Earlier quoted context omitted.
> Github is ripe for disruption and I hope it is disrupted soon. It's an expensive, low revenue generating site. There are, and have always been, competitors, including "host it all yourself" solutions, but nothing has really stuck. How is it "ripe" for disruption?
They had $1b revenue in 2023 and now probably more than $2b in revenue... do you have cost figures showing what their expenses are?
That can't be cheap.
Re: The August 17 outage
#34"Since April, monthly commits have grown from 1.4 billion to 2.9 billion. " Wow, that is some incredible growth in a really short time.
Re: The August 17 outage
#35Github down, no hard drives available, no memory available, thanks AI! Seems like we are headed for Tech Gridlock.
What they can implement is to slowdown the commit rate, rate limt or just queue-up messages not to overburden their downstream service. I don't think GH has any of those, but just keep scaling, but that scaling failed. Just bad architectural decisions from the postmortem. -- It will only get worse due to AIs spawning massive commits, and they don't have unlimited cloud resource. They can scale but not scalable in ter…
> The immediate cause of the failure was network saturation on load balancers in Central US due to a new peak in traffic. Originally this was caused by an Istio sidecar pod reaching its concurrency limits and failing to auto scale correctly because of a misconfigured policy that watched host service but not sidecar limits. One failure cascaded to more and eventually four HAProxy nodes exhausted their flow limits, degrading the gateway auth path and causing widespread authentication latency and failures. The problem was worsened by optimistic retry logic which overloaded internal load balancers. Pausing HAProxy on those nodes simultaneously produced immediate broad recovery.
Re: The August 17 outage
#36Great read - I'm glad they realize there's work ahead but what I'm missing is: * Paid customers: we know you pay us often a ton of money, and we burn your month on actions during these outages - we'll refund you for the days we spent your money and gave you no value. * Paid customer: We know you put your trust in us, so we'll ensure we have a separate pool of capacity to ensure we can keep that trust. * Paid customer…
Re: The August 17 outage
#37Earlier quoted context omitted.
Baseless accusation made from a position of zero information.
Opinion based on stated facts. Please share the information you have which contradicts the conclusions I have drawn from Github's statement. (And we know they're liars. They report very few of the actual incidents they have; see for example https://mrshu.github.io/github-statuses/ )
Re: The August 17 outage
#38Earlier quoted context omitted.
> We installed as much hardware as available power allowed in our existing data centers while accelerating our migration to Azure. And from the RCA [1]: > The immediate cause of the failure was network saturation on load balancers in Central US due to a new peak in traffic. [1]: https://www.githubstatus.com/incidents/zkxwbgr0cnmx
"While accelerating our migration to Azure," meaning, they will only solve problems if it helps them also use Azure more. It is unbelivable that aload of 2.8b commits was totally fine, and a load of 2.9b was a sitewide outage, unless they have no reporting or their tooling is completely incompetent. If things can fall apart so easily, throwing more capacity at the problem won't fix it.
These backlogs can cause clients to make more retries, exacerbating the problem. Potentially further cascading through the system.
Re: The August 17 outage
#39"We are committed to fixing these problems, as long as it doesn't involve buying things other than AI computers, hiring humans, or using non-Microsoft products." Calling Azure the solution to this problem when it is in fact the source of most of these problems is just fantastic doublespeak. Github is ripe for disruption and I hope it is disrupted soon.
Re: The August 17 outage
#40Earlier quoted context omitted.
> Github is ripe for disruption and I hope it is disrupted soon. It's an expensive, low revenue generating site. There are, and have always been, competitors, including "host it all yourself" solutions, but nothing has really stuck. How is it "ripe" for disruption?
They had $1b revenue in 2023 and now probably more than $2b in revenue... do you have cost figures showing what their expenses are?
Compute and Storage for Free Tiers: Hosting code for over 150 million developers and processing over 2 billion GitHub Actions (CI/CD) workflows a month requires astronomical server power and data storage. The "Free" tier is a massive cost sink that Microsoft treats as a loss-leader marketing expense
Let me know when you understand how that's not free.