> We noticed a strong correlation between crazy utilization spikes and CI failure rates. This is interesting, and is something I've also suspected on many CI systems that offer free public runners (CircleCI, GitHub Actions, etc.). For seemingly no reason at all, tests were very flaky and unstable in CI, which couldn't be reproduced on local machines. I tried everything from resource-limited containers, to identically…
That was one of the reasons we ended up setting up our own runners. Didn't mention in the post but we use spot VM instances.