Live data from Hacker News

ECC and DDR5

etbe.coker.com.au

11–20 of 162 posts

Re: ECC and DDR5

#11

"Ideally we would have some government action to force this given the ongoing cost to society in corrupted data" Why is it so often Australia with the "We must force unnecessary things upon people" attitude?

Because despite geographical location, Australia is a western country, and all western countries are like that.

Re: ECC and DDR5

#12

Interesting article, and I agree that ECC needs to be more widely available. But I think this is a major overreach: > We need ECC RAM to be more widely used. Ideally we would have some government action to force this given the ongoing cost to society in corrupted data and lost time due to RAM hardware errors. I think that at minimum we need sufficient taxes on non-ECC RAM (and EC4 RAM for DDR5) to make it more expens…

Important data should probably be worked on and stored with error-correcting codes, in addition to the hardware-provided ones. The Btrfs mention in TFA suggests that maintaining integrity of the Btrfs metadata is not considered part of the job description of Btrfs, but perhaps it should be.

Re: ECC and DDR5

#13

"Ideally we would have some government action to force this given the ongoing cost to society in corrupted data" Why is it so often Australia with the "We must force unnecessary things upon people" attitude?

That particular quote did stand out to me. I think the counter claim is that if everyone uses software everywhere for everything, there is a very real societal cost for damaged data and if it's hard to detect then there is no way to prevent or measure harm. I think the "invisible hand" free market is showing some serious strain in modern times. People stopped using leaded-gas because it was causing measurable decreases in IQ for generations. I don't care how much you want to save a few dollars if it harms a lot of people and is difficult to measure. Everything else seems like a race to the bottom price so pushing the market in this manner doesn't bother me in the least. If you are running a server with others' data at stake, it seems like an obvious benefit for society on the whole to address the problem rather than serve some cheap datacenter operators.

Re: ECC and DDR5

#14

"Ideally we would have some government action to force this given the ongoing cost to society in corrupted data" Why is it so often Australia with the "We must force unnecessary things upon people" attitude?

I think it's a pattern that reinforces itself, as a nanny state develops it's opinion of it's people worsens and it's more willing to use force to compel them.

Re: ECC and DDR5

#15
> One thing that concerns me is the possibility of on-die ECC interacting with ECC on the motherboard and reducing it’s effectiveness

DDR5 on-die ECC detects and corrects one-bit errors. It cannot detect two-bit errors, so it will miscorrect some of them into three-bit errors. My understanding is that the on-die error correction scheme is specifically specially designed such that the resulting three-bit errors are mathematically guaranteed to be detected as uncorrectable two-bit errors by a standard full system-level ECC running on top of the on-die ECC. But I've never found a real authoritative reference that directly says that.

Re: ECC and DDR5

#16

The problem I have with ECC discussions is that it’s always anecdotal. There is one Google study finding memory flips with some probability in their datacenters, but it was awhile ago. Everyone repeats the statement DDR5 on-die ECC isn’t a replacement for regular ECC, but nobody ever shows data on DDR5 bitflips. That’s not a knock on this article per se, but it’d be nice to have actual data here.

>> Everyone repeats the statement DDR5 on-die ECC isn’t a replacement for regular ECC, but nobody ever shows data on DDR5 bitflips. That's because the on-die ECC cannot detect errors in the transmission of data from the DIMMs to the CPU. It seems like on-die ECC is meant to "hide" some level of errors to make less reliable RAM chips appear good. I have no idea if that's a reasonable way to do it, but it seems like it…

Right. Everyone agrees that real ECC is more robust, and production servers of course run it. But it seems to me that if most RAM bitflips are in the module itself (rather than in transmission to CPU), then on-die ECC could in theory make the probability of a bitflip in practice so low that for things like home servers it isn’t worth the price premium for real ECC. But without data on this, everyone is just speculating.

Re: ECC and DDR5

#17
post #6

Earlier quoted context omitted.

> ECC RAM is not needed for every application and mandating it for every gaming rig and machine used to run Chrome will just increase prices unnecessarily. The price increase will be very small if manufacturers are no longer able to segment the market to make a higher margin on those who need it. No one wants even games or web browsers to crash randomly. > If they want to eat that cost, it should be up to them. They…

Without quantifying the cost to society that someone's web-browser crashes sometimes, it's hard to make a rational case that this issue is so dire that it needs specific regulation. I'm not sure making it illegal to sell consumer grade hardware is the boon for the people you imagine

Crash is the best-case outcome for a browser bit flip. Here's bitsquatting:

https://ripe92.ripe.net/programme/meeting-plan/sessions/112/...

Re: ECC and DDR5

#18

Earlier quoted context omitted.

>> Everyone repeats the statement DDR5 on-die ECC isn’t a replacement for regular ECC, but nobody ever shows data on DDR5 bitflips. That's because the on-die ECC cannot detect errors in the transmission of data from the DIMMs to the CPU. It seems like on-die ECC is meant to "hide" some level of errors to make less reliable RAM chips appear good. I have no idea if that's a reasonable way to do it, but it seems like it…

Right. Everyone agrees that real ECC is more robust, and production servers of course run it. But it seems to me that if most RAM bitflips are in the module itself (rather than in transmission to CPU), then on-die ECC could in theory make the probability of a bitflip in practice so low that for things like home servers it isn’t worth the price premium for real ECC. But without data on this, everyone is just speculati…

This is the attitude that needs to change. If RAM is 25% of the system cost (say for a home server in your example), then the cost for ECC is 3% of the system cost. There is no world where accepting data corruption in order to save 3% cost is a reasonable tradeoff. ECC should be standard, period.

Re: ECC and DDR5

#19
post #5

Got ECC udimm for my ddr4 server and was surprised to see it picking up errors occasionally (once every couple months). Likely from a weakness in one of the sticks. This far it’s always corrected it though so opted to keep them anyway (nobody wants to be minus 32gb in these trying memory times) With normal sticks I’d not have know that there is a potential issue.

You should be able to get the stick replaced under warranty. Every memory manufacturer I have ever seen has a lifetime warranty. As long as you RMA it before replacements become unavailable.

Re: ECC and DDR5

#20

The problem I have with ECC discussions is that it’s always anecdotal. There is one Google study finding memory flips with some probability in their datacenters, but it was awhile ago. Everyone repeats the statement DDR5 on-die ECC isn’t a replacement for regular ECC, but nobody ever shows data on DDR5 bitflips. That’s not a knock on this article per se, but it’d be nice to have actual data here.

As TFA says, it is very likely that the hyperscalers continue to make studies about DRAM reliability in their servers, but they do not publish them, either because they believe that such data could be useful for competitors, or, more likely, they might have NDAs with the memory vendors.

When you do not own a great number of servers, you cannot provide anything else except anecdotal evidence.

Because the vast majority of PCs do not have ECC memory, even the companies that have a great number of PCs have no idea about the reliability of the memory used in those PCs, because it is very difficult to distinguish memory defects from the huge number of software bugs.

Moreover, the frequency of memory errors is proportional with the amount of memory you have, so those with tiny amounts of memory, like 8 GB, are much less likely to encounter memory errors than those who have 32 GB or 64 GB of DRAM in their PCs.

I have tried to always use ECC memory in my computers, whenever possible. Even my laptop is an older Dell Precision, with ECC memory.

But I can also offer only anecdotal testimony.

I have also seen cases like that mentioned in TFA, where the existence of ECC allowed to discover that some modules were not seated well in their sockets, so reseating them prevented the reappearance of errors.

I have also seen a case when a laptop was not used for a long time and it was stored in a rather humid place, so the contacts in the SODIMM sockets had oxidized, which resulted in frequent memory errors. Scrubbing vigorously the contacts and reinserting the memory modules solved the problem.

I have also seen many cases where certain memory modules degraded after many years of use, e.g. 5 or more years of 24/7 use, and they began to have frequent errors, e.g. multiple errors per day, even if when they were new the error rate could have been of one error per year or even less.

In such cases ECC was extremely useful, because the DIMM that had become worn out could be identified and replaced and the server could work fine some more years.

Post reply on HN