Live data from Hacker News

ECC and DDR5

etbe.coker.com.au

71–80 of 162 posts

Re: ECC and DDR5

#71
post #22

Earlier quoted context omitted.

My Linux installation came with an AMD MCE driver that reports them to dmesg. It logs a few corrected errors a day.

A few errors per day seems much too frequent, unless you live at a high altitude. Normal good DIMMs, at least when new, should not have errors more frequently than one error per many months. Frequent errors may appear with memories that are not seated well in their sockets, or which are old, at least several years old. Frequent errors may also be caused by more general computer problems, like a bad power supply unit.

Are DIMMs (a.k.a. UDIMM) still a thing?

I would think most servers and workstations would be RDIMM (Registered DIMM) by now and consumer stuff uses soldered down memory. Memory failing because it's old is definitely a thing, and very possible in this scenario, but I feel like I haven't seen errors due to physical insertion, that were not caught immediately by POST, in years.

But maybe it's just me. Happy to be lucky. :)

Re: ECC and DDR5

#72
post #51

I've always thought it odd that we don't use ECC as a standard on all computers. I think people really downplay the impact of memory issues. They can be devastating, especially over time. They will slowly corrupt your file system, documents, binary files, code, everything. You notice when things like your compressed files start giving CRC/checksum errors or your downloads don't match SHA-512. You are then left with t…

This is/was classic market segmentation. Want ECC? Pay up for Xeon (server/workstation CPU)

Re: ECC and DDR5

#73
post #50

Earlier quoted context omitted.

The primary source of uncorrelated bit-flips is, effectively, cosmic radiation. (Correlated bit-flips are likely from a manufacturing error... or, at least in one case, from excessively radioactive ceramic packaging.) The more atmosphere you have to randomly absorb high-energy photons, the better. Higher density memory drops fewer electrons in each well, which means a lower-energy photon can change the state. Corolar…

>The primary source of uncorrelated bit-flips is, effectively, cosmic radiation. This gets said a lot but with little evidence. The mario speedrun bit flip has been replicated with marginal connector insertion. https://youtu.be/vj8DzA9y8ls

There is a lot of direct evidence from increased error rates in computers on airplanes.

Re: ECC and DDR5

#74
post #50

Earlier quoted context omitted.

Curious to learn more here - why would altitude play a role here? And when you say high altitude - are we talking La Paz, Bolivia (~12k feet) or Denver, CO (~5k feet)?

The primary source of uncorrelated bit-flips is, effectively, cosmic radiation. (Correlated bit-flips are likely from a manufacturing error... or, at least in one case, from excessively radioactive ceramic packaging.) The more atmosphere you have to randomly absorb high-energy photons, the better. Higher density memory drops fewer electrons in each well, which means a lower-energy photon can change the state. Corolar…

FYI, it's not so much photons (gamma rays) as muons.

Muons are nasty little buggers.

Re: ECC and DDR5

#75

Earlier quoted context omitted.

The mechanical issues also act through increasing the sensitivity to electrical noise. The imperfect contacts are equivalent to adding series resistors, possibly in parallel with parasitic capacitors, on the link traces, allowing a greater amplitude for the noise pulses.

In addition to reducing the energy content of the actual desired signal.

Correct.

Re: ECC and DDR5

#76

Earlier quoted context omitted.

>The primary source of uncorrelated bit-flips is, effectively, cosmic radiation. This gets said a lot but with little evidence. The mario speedrun bit flip has been replicated with marginal connector insertion. https://youtu.be/vj8DzA9y8ls

There is a lot of direct evidence from increased error rates in computers on airplanes.

That's a much weaker statement than "the primary source of uncorrelated bit flips is cosmic radiation".

Re: ECC and DDR5

#77
post #57

Earlier quoted context omitted.

The motherboard manufacturers test their motherboards, to ensure that they typically work with overclocked memory modules. If you buy a complete gaming computer which includes overclocked memory, then hopefully the vendor has done a burn-in and has tested for some time the computer, as sold. However, if you buy the CPU yourself, neither Intel nor AMD provides any guarantee that the CPU will work with overclocked memo…

> However, if you buy the CPU yourself, neither Intel nor AMD provides any guarantee that the CPU will work with overclocked memory. There's no guarantee that it works with standard clocked memory either.

Both AMD and Intel publish on their site, for each CPU model, the maximum supported memory speed.

If the CPU does not work at that speed, you are entitled to a replacement or a refund.

For instance, AMD Ryzen™ 9 9950X3D is guaranteed to work with DDR5-5600, but with nothing faster. No AMD non-server CPU goes above this limit for DDR5.

Some of the Intel CPUs are guaranteed to work with faster memories, e.g. Core Ultra 9 285K with up to DDR5-6400, and the very recently launched Core Ultra 7 270K PLUS with up to DDR5-7200.

Re: ECC and DDR5

#78

Earlier quoted context omitted.

A few errors per day seems much too frequent, unless you live at a high altitude. Normal good DIMMs, at least when new, should not have errors more frequently than one error per many months. Frequent errors may appear with memories that are not seated well in their sockets, or which are old, at least several years old. Frequent errors may also be caused by more general computer problems, like a bad power supply unit.

Are DIMMs (a.k.a. UDIMM) still a thing? I would think most servers and workstations would be RDIMM (Registered DIMM) by now and consumer stuff uses soldered down memory. Memory failing because it's old is definitely a thing, and very possible in this scenario, but I feel like I haven't seen errors due to physical insertion, that were not caught immediately by POST, in years. But maybe it's just me. Happy to be lucky.…

The client/consumer desktop market hasn't actually disappeared, and soldered memory is almost unheard-of in that market segment.

Re: ECC and DDR5

#79
post #36

Earlier quoted context omitted.

> There is no world where accepting data corruption in order to save 3% cost is a reasonable tradeoff. It is if you're doing something where corruption is detectable after the fact, and happens rarely enough that redoing affected work adds less than 3% overhead.

If you don't have ECC, you don't even know that your CPU is executing the correct code. There goes the "detectable" part.

DRAM is just one level in the memory hierarchy and it can absolutely be tested by the other levels in the system. It won't be continually monitored, and there will be a performance penalty, and it's a PITA, but it can definitely be done without hardware ECC. Even some of that can be minimized as most tests are statistical in nature anyway.

Re: ECC and DDR5

#80
post #39
post #18

Earlier quoted context omitted.

This is the attitude that needs to change. If RAM is 25% of the system cost (say for a home server in your example), then the cost for ECC is 3% of the system cost. There is no world where accepting data corruption in order to save 3% cost is a reasonable tradeoff. ECC should be standard, period.

Let me kmow where I can get ECC RAM that's only 12% more than the price of non-ECC and I'll strongly consider it the next time I do a build. (Otoh, prices are high, so I'm not building anything for a while unless stuff breaks)

Some years ago, it was easy to find ECC modules with a so small price difference.

From that time, I have a few old computers with 64 GB or 128 GB of DDR4 ECC memory.

Unfortunately, after the passage to DDR5, DDR5 ECC UDIMMs have become hard to find, and even when you could find them, the price difference could be as high as 30% to 50%.

Post reply on HN