I struggle to find the correct descriptor for a counter-example, wherein You Really Want Success for the process as a whole, but it is acceptable for a sliver of it to fail, in the the context of ETL. I have an ETL I am told (I switched jobs) that is still working, from 2008. It was built to be a tank, and I also did another forbidden thing: Pokemon Exception Handling. It's a guideline, not a law of physics, and it i…
Funny thing, I had same experience. I also built robust ETL. It was for ingesting and manipulating financial data from 30+ different banks and I also did Pokemon Exception Handling to make it robust. In general my philosophy is: I don't want to wake up at 4am unless it's urgent. What can I do to gracefully handle failures to achieve that goal?
My feeling is that, if you have a foundational business process like this, it should be designed to be maintained if there is a serious problem, and it ought to keep working. I know, haha, "the Internet is a series of tubes" but I really wanted this thing to be like a chunk of very uninteresting ductwork that just moves air from one place to another: it should just do its job with as much fanfare.