Live data from Hacker News

Imaging a Hard Drive with Non-ECC Memory – What Could Go Wrong?

blog.robertelder.org

31–40 of 79 posts

Re: Imaging a Hard Drive with Non-ECC Memory – What Could Go Wrong?

#31
post #7

I've had a lot of really strange bugs and data loss with my current build (Ryzen with Gskill memory). After running a memtest for 24h i finally saw that two of the four ram sticks were faulty (two bit flips on each only rarely and on a specific test). The company changed them but now a year later without any issues I have another one that failed in exactly the same way. This is the last time I build a non-ECC system…

You might want to take a look at your PSU. That seems like a suspicious amount of RAM failure to me. How old is it and what model, if you don't mind me asking?

PSU is something I never cheap out on. Always pays for itself in the end. A bad PSU can kill your whole system.

Re: Imaging a Hard Drive with Non-ECC Memory – What Could Go Wrong?

#32
post #7

I've had a lot of really strange bugs and data loss with my current build (Ryzen with Gskill memory). After running a memtest for 24h i finally saw that two of the four ram sticks were faulty (two bit flips on each only rarely and on a specific test). The company changed them but now a year later without any issues I have another one that failed in exactly the same way. This is the last time I build a non-ECC system…

You might want to take a look at your PSU. That seems like a suspicious amount of RAM failure to me. How old is it and what model, if you don't mind me asking? PSU is something I never cheap out on. Always pays for itself in the end. A bad PSU can kill your whole system.

5 years. and the current ram sticks (the 3 that work) can sustain a 3 day memtest with 0 errors (at 3600). They were all able to handle that when I received them back from warranty.

Re: Imaging a Hard Drive with Non-ECC Memory – What Could Go Wrong?

#34
post #5

I ran a cluster of ~30k blade based computers booting entirely off iPXE. They didn't have any onboard ssd/disk storage or ECC memory. Every day, a few of them would randomly lock up, they'd reboot with a fresh network image and keep on humming.

> Every day, a few of them would randomly lock up, they'd reboot with a fresh network image and keep on humming. There same ones, or random new machines every time?

Totally random.

Re: Imaging a Hard Drive with Non-ECC Memory – What Could Go Wrong?

#35
post #29
post #5

I ran a cluster of ~30k blade based computers booting entirely off iPXE. They didn't have any onboard ssd/disk storage or ECC memory. Every day, a few of them would randomly lock up, they'd reboot with a fresh network image and keep on humming.

How do you even get that many computers without ECC? I think all the blades I've seen have ECC as baseline spec.

They were purpose built and didn't require 100% uptime.

Re: Imaging a Hard Drive with Non-ECC Memory – What Could Go Wrong?

#36
post #4

ECC is good, and I genuinely wish it were more common. Thankfully, Ryzen CPUs support ECC by default (except for pre-7000 series with integrated graphics that aren't "Pro" versions), so long as the motherboard does, too (like all ASRock that I've seen). I'm running several Ryzen servers with ECC. On the other hand, there are many, many systems out there that don't have ECC, nor do they have the option to have ECC. Wh…

ECC is now baked into the DDR5 spec. Great news!

Re: Imaging a Hard Drive with Non-ECC Memory – What Could Go Wrong?

#37
post #4

ECC is good, and I genuinely wish it were more common. Thankfully, Ryzen CPUs support ECC by default (except for pre-7000 series with integrated graphics that aren't "Pro" versions), so long as the motherboard does, too (like all ASRock that I've seen). I'm running several Ryzen servers with ECC. On the other hand, there are many, many systems out there that don't have ECC, nor do they have the option to have ECC. Wh…

Intel finally started supporting ECC again on their consumer CPUs, but if only the W680 didn't cost ~$400 starting.

Re: Imaging a Hard Drive with Non-ECC Memory – What Could Go Wrong?

#38
>Does increased heat increase the likelihood of memory errors? I think it does.

I just got through a round of overclocking my memory. Yes, heat does.

>tRFC is the number of cycles for which the DRAM capacitors are "recharged" or refreshed. Because capacitor charge loss is proportional to temperature, RAM operating at higher temperatures may need substantially higher tRFC values.

https://github.com/integralfx/MemTestHelper/blob/oc-guide/DD...

Re: Imaging a Hard Drive with Non-ECC Memory – What Could Go Wrong?

#39
post #32

Earlier quoted context omitted.

You might want to take a look at your PSU. That seems like a suspicious amount of RAM failure to me. How old is it and what model, if you don't mind me asking? PSU is something I never cheap out on. Always pays for itself in the end. A bad PSU can kill your whole system.

5 years. and the current ram sticks (the 3 that work) can sustain a 3 day memtest with 0 errors (at 3600). They were all able to handle that when I received them back from warranty.

5 years could be getting up there if it's a mid range to low end PSU in terms of reliability. You might want to see if you find your PSU in this list:

https://cultists.network/140/psu-tier-list/

Re: Imaging a Hard Drive with Non-ECC Memory – What Could Go Wrong?

#40

Earlier quoted context omitted.

> Every day, a few of them would randomly lock up, they'd reboot with a fresh network image and keep on humming. There same ones, or random new machines every time?

Totally random.

Could easily be software or some other marginal hardware bug though.
Post reply on HN