Earlier quoted context omitted.
Update 16:36 UTC - We identified the problematic component and have taken corrective actions. There are strong signs of recovery but we are still working to completely restore service, with error rates still remaining slightly elevated. We will post further updates as recovery continues. just shy of 3 hours to find the issue
Furiously prompting Claude, "please, please fix this!!!!"
Incident with Github.com [resolved]
701–710 of 1001 posts
Re: Incident with Github.com [resolved]
#702To everyone who is angry: calm down. Github’s servers are constantly on fire as their usage increased something like 50x due to LLM sloppers pushing large amounts of trash code Unless you have a emergency hotfix (you don’t), go hit the gym or walk outside. If a work tool going down triggers you enough mentally to start angrily ranting online, it’s a sign you need to chill out and focus more on your health
They get paid millions of dollars by my company. They need to get it together.
Re: Incident with Github.com [resolved]
#703Earlier quoted context omitted.
Weird, I've been told you can just throw AI at all your logs and metrics and it can magically solve all the things...
Maybe they are dedicated hand coders!
So a "component" failed? That sounds like a software component. We just saw in https://www.wiz.io/blog/red-agent-snowflake-copilot-cicd-bug how sloppy bot coding is.
Re: Incident with Github.com [resolved]
#704Earlier quoted context omitted.
Scaling can take time but they've had at least a year to prepare. It's not like any demand increase they've seen was overnight... At most normal companies you monitor your systems and address potential bottlenecks before they reach a tipping point. And generally you want enough headroom that a sudden 2-3x increase in demand wouldn't take out the service. Either Github's technical leadership/talent is completely out o…
More likely, as seen in a lot of companies, perhaps infra teams got whittled down (or frozen HC, less than BAU, etc) with resources reallocated to AI org units.
Re: Incident with Github.com [resolved]
#705Earlier quoted context omitted.
Because any price at all will immediately cause users to shift to another platform, and GitHub's value is that it is _the_ place to put your code on the internet.
The network effects isn’t that much, and like what we see with Steam, you want to be where the audience is. More realistically, MAU or something is a metric/OKR. It’s easy to imagine why a business wouldn’t want to cap that.
The audience will go away to somewhere you aren't. Then what?
Re: Incident with Github.com [resolved]
#706I had a lot of goodwill for GitHub but I think today is the tipping point. Looking at a unicorn page, I feel this lingering hope that it's transient (like it usually was in the old days) but my mind reassures me it's probably going to be a long full outage again. The hope is dead.
I’m willing to give them a break as I’m assuming they have a lot of scaling problems due to the influx of LLM assisted coding. But maybe I’m wrong?
It has really just coasted on GH's network effect because everybody uses GH for git by default. It doesn't really have much direct competition and anything that does offer something like GH just doesn't have the same mindshare.
In that way, GH became more like MS.
Re: Incident with Github.com [resolved]
#707Earlier quoted context omitted.
People overestimate how much they care about stuff. Moving off GitHub would be more costly for us than having 5% downtime. Obviously there’s a tipping point, but it shows people are tolerant given the price tags.
For a reasonably common class of deployment and organization (E.g. an online service where the cost of a short outage is significant) it becomes a pain point. For instance if there's a sudden security flap and you need to re-spin your system and redeploy, but that process is gated on GitHub working, now you're screwed if that security flap happens when GitHub is in its 5% down time. Basically you've coupled your upti…
Re: Incident with Github.com [resolved]
#708Earlier quoted context omitted.
Doubly so given this is hardly a surprise. We've been on this trajectory for at least a couple of years now. They don't get to shrug, point at 10x volume, and act like they've been blindsided.
They had large increase in volume, and trying to move to Azure at the same time. I don't envy them for either work they need to do. But also don't feel pithy because it's Microslop, at the end of the day.
Managers and executives though? Those I do blame. Surely at this point it should be blindingly obvious their current strategy is not working.
Re: Incident with Github.com [resolved]
#709Cannot even merge code, this is the limit for me. We are prioritising migration to a different platform and this time going to decouple CI - right now we have too many points of failure on one vendor.
Apologies if you already tried this, but I have been able to merge and get some things done via command line. e.g.: gh pr merge 5062 -R --squash --delete-branch
Re: Incident with Github.com [resolved]
#710Earlier quoted context omitted.
How many weeks/months can they use this excuse? They are literally are at the forefront of this emerging industry and are capturing untold value. To let their product suffer and potentially lose market share because of it is extremely foolish
Scaling can take time but they've had at least a year to prepare. It's not like any demand increase they've seen was overnight... At most normal companies you monitor your systems and address potential bottlenecks before they reach a tipping point. And generally you want enough headroom that a sudden 2-3x increase in demand wouldn't take out the service. Either Github's technical leadership/talent is completely out o…