.. and that's how people end up with ZFS. My story was cheap SATA cables (or was it cheap power supplies? I never found out) introducing silent corruption and breaking decades of archived files.
It’s pretty remarkable actually just how bad that failure mode was. How did mdadm manage to cause such havoc after the broken device was reattached?
And the drive would return bad data instead of just erroring out - this is sometimes a difference between “home use” drives vs enterprise - enterprise assume you’re in a multi redundant raid setup and instead of retrying will just fail out fast and let the raid take care of it. Home drives will continue retrying and perhaps even eventually give you what they got - the idea being you’d rather have one byte bad of a word doc than lose a 4K chunk of it.
ZFS catches these tricks because the checksums will fail. Multi-redundancy RAID can do it - but since most are enterprise they are built with the assumption the drive will return good data - or none.
This can be the worst scenario as corruption can be silent and infect all backups after a time.