Live data from Hacker News

My 71 TiB ZFS NAS After 10 Years and Zero Drive Failures

louwrentius.com

21–30 of 314 posts

Re: My 71 TiB ZFS NAS After 10 Years and Zero Drive Failures

#22
I have a similar-sized array which I also only power on nightly to receive backups, or occasionally when I need access to it for a week or two at a time.

It's a whitebox RAID6 running NTFS (tried ReFS, didn't like it), and has been around for 12+ years, although I've upgraded the drives a couple times (2TB --> 4TB --> 16TB) - the older Areca RAID controllers make it super simple to do this. Tools like Hard Disk Sentinel are awesome as well, to help catch drives before they fail.

I have an additional, smaller array that runs 24x7, which has been through similar upgrade cycles, plus a handful of clients with whitebox storage arrays that have lasted over a decade. Usually the client ones are more abused (poor temperature control when they delay fixing their serveroom A/C for months but keep cramming in new heat-generating equipment, UPS batteries not replaced diligently after staff turnover, etc...).

Do I notice a difference in drive lifespan between the ones that are mostly-off vs. the ones that are always-on? Hard to say. It's too small a sample size and possibly too much variance in 'abuse' between them. But definitely seen a failure rate differential between the ones that have been maintained and kept cool, vs. allowed to get hotter than is healthy.

I can attest those 4TB HGST drives mentioned in the article were tanks. Anecdotally, they're the most reliable ones I've ever owned. And I have a more reasonable sample size there as I was buying dozens at a time for various clients back in the day.

Re: My 71 TiB ZFS NAS After 10 Years and Zero Drive Failures

#23
post #9

the 'secret' is not that you turn them off. it's simply luck. I have 4TB HGST drives running 24/7 for over a decade. ok, not 24 but 8, and also 0 failures. But I'm also lucky, like you. Some of the people I know have several RMAs with the same drives so there's that. My main question is: What is it that takes 71TB but can be turned off most of the time? Is this the server you store backups?

It can be luck, but with 24 drives, it feels very lucky. Somebody with proper statistics knowledge can probably calculate the risk with a guestimated 1% yearly failure rate how likely it would be to have all 24 drives remaining. And remember, my previous NAS with 20 drives also didn't have any failures. So N=44, how lucky must I be? It's for residential usage, and if I need some data, I often just copy it over 10Gbit…

The failure rate is not truly random with a nice normal distribution of failures over time. There are sometimes higher rates in specific batches or they can start failing altogether etc.

Backblaze reports always are interesting insights into how consumers drives behave under constant load.

Re: My 71 TiB ZFS NAS After 10 Years and Zero Drive Failures

#24
post #6

What does one do with all this storage?

This is a drop in the bucket for photographers, videographers, and general backups of RAW / high resolution videos from mobile devices. 80TB [usable] was "just enough" for my household in 2016.

Exactly that. I'm not even shooting in ProRes or similar "raw" video. But one video project easily takes 3TB. And I'm not even a professional.

Re: My 71 TiB ZFS NAS After 10 Years and Zero Drive Failures

#25
post #2

There have been drives where power cycling was hazardous. So, whilst I agree to the model, it shouldn't be assumed this is always good, all the time, for all people. Some SSD need to be powered periodically. The duty cycle for a NAS probably meets that burden. Probably good, definitely cheaper power costs. Those extra grease on the axle drives were a blip in time. I wonder if backblaze do a drive on-off lifetime stat…

> There have been drives where power cycling was hazardous. I know about this story from 30+ years ago. It may have been true then. It may be even true now. Yet, in my case, I don't power cycle these drives often. At most a few times a month. I can't say or prove it's a huge risk. I only believe it's not. I have accepted this risk for over 15+ years. Update: remember that hard drives have an option to spin down when…

I debated posting because it felt like shitstirring. I think overwhelmingly what you're doing is right. And if a remote power on eg WOL works on the device, so much the better. If I could wish for one thing, it's mods to code or documentation of how to handle drive power down on zfs. The rumour mill is zfs doesn't like spindown.

Re: My 71 TiB ZFS NAS After 10 Years and Zero Drive Failures

#26

Earlier quoted context omitted.

It can be luck, but with 24 drives, it feels very lucky. Somebody with proper statistics knowledge can probably calculate the risk with a guestimated 1% yearly failure rate how likely it would be to have all 24 drives remaining. And remember, my previous NAS with 20 drives also didn't have any failures. So N=44, how lucky must I be? It's for residential usage, and if I need some data, I often just copy it over 10Gbit…

We don't really have to guess. Backblaze posted their stats for 4 TB HGST drives for 2024, and of their 10,000 drives, 5 failed. If OP's 2014 4 TB HGST drives are anything like this, then this is just snake oil and magic rituals and it doesn't really matter what you do.

5 drives failed in Q1 2024. 8 died in Q2. That is still a very low failure rate.

Re: My 71 TiB ZFS NAS After 10 Years and Zero Drive Failures

#27
I have a mini PC + 4x external HDDs (I always bought used) on Windows 10 with ReFS since probably 2016 (recently upgraded to Win 11), maybe earlier. I don't bother powering off.

The only time I had problems is when I tried to add a 5th disk using a USB hub, which caused drives attached to the hub get disconnected randomly under load. This actually happened with 3 different hubs, so I since stopped trying to expand that monstrosity and just replace drives with larger ones instead. Don't use hubs for storage, majority of them are shitty.

Currently ~64TiB (less with redundancy).

Same as OP. No data loss, no broken drives.

A couple of years ago I also added an off-site 46TiB system with similar software, but a regular ATX with 3 or 4 internal drives because the spiderweb of mini PC + dangling USBs + power supplies for HDDs is too annoying.

I do weekly scrubs.

Some notes: https://lostmsu.github.io/ReFS/

Re: My 71 TiB ZFS NAS After 10 Years and Zero Drive Failures

#28

Earlier quoted context omitted.

It can be luck, but with 24 drives, it feels very lucky. Somebody with proper statistics knowledge can probably calculate the risk with a guestimated 1% yearly failure rate how likely it would be to have all 24 drives remaining. And remember, my previous NAS with 20 drives also didn't have any failures. So N=44, how lucky must I be? It's for residential usage, and if I need some data, I often just copy it over 10Gbit…

We don't really have to guess. Backblaze posted their stats for 4 TB HGST drives for 2024, and of their 10,000 drives, 5 failed. If OP's 2014 4 TB HGST drives are anything like this, then this is just snake oil and magic rituals and it doesn't really matter what you do.

> If OP's 2014 4 TB HGST drives are anything like this, then this is just snake oil and magic rituals and it doesn't really matter what you do.

It might matter what you do, but we only have public data for people in datacenters. Not a whole lot of people with 10,000 drives are going to have them mostly turned off, and none of them shared their data.

Re: My 71 TiB ZFS NAS After 10 Years and Zero Drive Failures

#29
post #6

What does one do with all this storage?

If you prefer to own media instead of streaming it, are into photography, video editing, 3D modelling, any AI-related stuff (models add up) or are a digital hoarder/archivist you blow through storage rather quickly. I'm sure there are some other hobbies that routinely work with large file sizes.

Storage is cheap enough that rather than deleting 1000's of photos and never be able to reclaim or look at them again I'd rather buy another drive. I'd rather have a RAW of an 8 year old photo that I overlooked and decide I really like and want to edit/work with than a 87kb resized and compressed JPG of the same file. Same for a mostly-edited 240GB video file. What if I want or need to make some changes to it in the future? May as well hold onto it than have to re-edit the video or re-shoot the video if the original footage was also deleted.

Content creators have deleted their content often enough that if I enjoyed a video and think future me might enjoy rewatching the video - I download it rather than trust that I can still watch it in the future. Sites have been taken offline frequently enough that I download things. News sites keep restructuring and breaking all their old article links so I download the articles locally. JP artists are notorious for deleting their entire accounts and restarting under a new alias that I routinely archive entire Pixiv/Twitter accounts if I like their art as there is no guarantee it will still be there to enjoy the next day.

It all adds up and I'm approaching 2 million well-organized and (mostly) tagged media files in my Hydrus client [0]. I have many scripts to automate downloading and tagging content for these purposes. I very, very rarely delete things. My most frequent reason for deleting anything is "found in higher quality" which conceptually isn't really deleting.

Until storage costs become unreasonable I don't see my habits changing anytime soon. On the contrary - storage keeps getting cheaper and cheaper and new formats keep getting created to encode data more and more efficiently.

[0] https://hydrusnetwork.github.io/hydrus/index.html

Re: My 71 TiB ZFS NAS After 10 Years and Zero Drive Failures

#30

Earlier quoted context omitted.

We don't really have to guess. Backblaze posted their stats for 4 TB HGST drives for 2024, and of their 10,000 drives, 5 failed. If OP's 2014 4 TB HGST drives are anything like this, then this is just snake oil and magic rituals and it doesn't really matter what you do.

5 drives failed in Q1 2024. 8 died in Q2. That is still a very low failure rate.

Drives have a bathtub curve, but if you want you can be conservative and estimate first year failure rates throughout. So that's p=5/10000 for drive failure. So chance of no-failure per year (because of our assumption) is 1-p. So, chance of no-failure per ten year is (1-p)^10 or about 99.5%
Post reply on HN