alt.sysadmin.recovery lives on! albeit in a web app. wonder if usenet is still alive...
https://groups.google.com/forum/#!forum/alt.sysadmin.recover...
11–20 of 102 posts
alt.sysadmin.recovery lives on! albeit in a web app. wonder if usenet is still alive...
https://groups.google.com/forum/#!forum/alt.sysadmin.recover...
The customer.io story seems like a great example of why NOT to use budget providers like OVH and Hetzner for mission-critical applications. You get what you pay for.
Entirely different problem from a provider that loses one internet connection and their other links can't keep up with traffic, but you can still have major problems even if you're spending thousands of dollars a month compared to hundreds.
My Devops horror stories, one sentence each: - Somebody deployed new features on a Friday at 5pm. - Fifteen hundred machines running mod_perl. - Supporting Oracle - TWICE. - It turns out your entire infrastructure is dependent on a single 8U Sun Solaris machine from 15 years ago, and nobody knows where it is. - Troubleshooting a bug in a site, view source.... and see SQL in the JS.
My Devops horror stories, one sentence each: - Somebody deployed new features on a Friday at 5pm. - Fifteen hundred machines running mod_perl. - Supporting Oracle - TWICE. - It turns out your entire infrastructure is dependent on a single 8U Sun Solaris machine from 15 years ago, and nobody knows where it is. - Troubleshooting a bug in a site, view source.... and see SQL in the JS.
The customer.io story seems like a great example of why NOT to use budget providers like OVH and Hetzner for mission-critical applications. You get what you pay for.
The customer.io story seems like a great example of why NOT to use budget providers like OVH and Hetzner for mission-critical applications. You get what you pay for.
At a previous job we hosted with {HAL} out of Atlanta. A NOC operator there saw/heard/smelled something that indicated to him that he should hit the Big Red Switch. So he did. This removed power to every machine in that part of the DC.
After management confirmed that there was no life-threatening emergency, they started bringing everything back up. Only to have machines start going down again 20 minutes later, as their local UPSes ran out of juice. Someone had to walk around to every cage and recycle them all manually.
The customer.io story seems like a great example of why NOT to use budget providers like OVH and Hetzner for mission-critical applications. You get what you pay for.
Echoing some of the other comments... the place i worked at was one of the 5 biggest Rackspace customer, and that didn't stop them from regularly cutting traffic or bringing down servers for hours in some case. You might get better uptime overall from reputable providers, but ultimately it's all about distributed application/service architecture.
Forgot about tmpwatch, a default entry in the RHEL cron table to clear out old temp files.
4AM the next morning, recursive deletion on anything wiuth a change time older than n days.