Live data from Hacker News

An Empirical Analysis of Hardware Failures on a Million Consumer PCs

research.microsoft.com

71–80 of 80 posts

Re: An Empirical Analysis of Hardware Failures on a Million Consumer PCs

#72
post #62

When Microsoft, Google, or some university publish analysis of hardware failures across large numbers of machines, they always anonymize hardware vendors ("vendor A", "vendor B"). I understand the reasons (not alienating your hardware vendors), but will there ever be a research group who will disclose vendor names? Heck, I would pay for this information.

This is dated, only applies to laptops, but it breaks down failures by manufacturer: http://www.electronicsweekly.com/blogs/engineering-design-pr...

Re: An Empirical Analysis of Hardware Failures on a Million Consumer PCs

#73
post #69
post #62

When Microsoft, Google, or some university publish analysis of hardware failures across large numbers of machines, they always anonymize hardware vendors ("vendor A", "vendor B"). I understand the reasons (not alienating your hardware vendors), but will there ever be a research group who will disclose vendor names? Heck, I would pay for this information.

IMO such information would do more harm than good. By the time they could gather statistics, that model would be obsolete so the stats wouldn't help you buy new equipment. But the kind of people who are still griping about the Deathstar would use the information to troll non-stop.

BUT if we had consistent poor performance for some vendor in a certain category, we could infer that their offerings for that category will also be poor in the future.

Re: An Empirical Analysis of Hardware Failures on a Million Consumer PCs

#74
post #59
post #13

Earlier quoted context omitted.

Funny, I suspect that SSDs have much higher real-world failure rates. (My personal, limited, anecdotal evidence is that my 64 GB Crucial M4 SSD lasted about a year as the root drive in a busy Linux desktop, while I have a stack of about a dozen hard drives that have been retired due to being too small or too slow while still working fine.) Lack of moving parts is great, but flash allows a finite number of write cycle…

How heavily were you using the laptop? Did it fail from running out of write cycles or something else? Some people over on xtremesystems have done Endurance testing, and the 64 GB m4 took over 700TB of writing to for failure to occur, and 172 TB to reduce the MWI to 0. In a little over a year I have only written 4.1TB to my SSD in my desktop. Write cycles are very unlikely to run out for me before I replace the drive…

I don't have actual numbers, but it's my primary desktop at home, and it saw everyday use.

Re: An Empirical Analysis of Hardware Failures on a Million Consumer PCs

#75
post #39

Earlier quoted context omitted.

Perhaps the difference is in testing. A large mfg, I would imagine, would test a configuration repeatedly before making it available, and then, once approved, the individual systems would go through burn-in, probably with more rigor than beige boxes. So even beige boxes with pricier (but unproven configurations) might suffer from grater failure rates. In addition, large mfgs might be able to demand better "lots" from…

Dell tests nothing. Parts in one door, assembly, shipped out other door to consumer. Essentially you the consumer are doing the burn-in. Its cheaper for Dell to replace failed machines. The cost to burn-in (and the time!) is large.

I was wrong! A former Dell employee tells of touring a plant and seeing the test station - hydra-like cables dangling from the ceiling with 1 plug for every hole in the computer. It would get plugged in, network-boot diagnostics and run for some time before being passed. But this was 12 years ago...

Re: An Empirical Analysis of Hardware Failures on a Million Consumer PCs

#76
post #71

I've always said that smart hardware tinkerers underclock . It produces less heat, and results in a quieter machine. I always suspected it improves reliability.

Or you could save money and just buy a lower bin.

Well, because heat dissipation is proportional to the square of the voltage, you end up giving up a little bit of performance but save a whole lot of heat. In experiential terms, you never miss performance but often notice a whole lot less fan noise.

Buying from a lower bin, you're getting a crappier processor, which might give you less latitude to save heat. This would probably be worth measuring and writing an article about. Also, I tend to buy lower clocked processors as it is.

Re: An Empirical Analysis of Hardware Failures on a Million Consumer PCs

#77
post #71

Earlier quoted context omitted.

Or you could save money and just buy a lower bin.

Well, because heat dissipation is proportional to the square of the voltage, you end up giving up a little bit of performance but save a whole lot of heat. In experiential terms, you never miss performance but often notice a whole lot less fan noise. Buying from a lower bin, you're getting a crappier processor, which might give you less latitude to save heat. This would probably be worth measuring and writing an arti…

I think it's likely that all lower bins are artificial, so e.g. underclocking a 2.4 GHz down to 2.0 is probably exactly the same as if you bought the 2.0. But yeah, it would be worth measuring.

Re: An Empirical Analysis of Hardware Failures on a Million Consumer PCs

#78
post #77

Earlier quoted context omitted.

Well, because heat dissipation is proportional to the square of the voltage, you end up giving up a little bit of performance but save a whole lot of heat. In experiential terms, you never miss performance but often notice a whole lot less fan noise. Buying from a lower bin, you're getting a crappier processor, which might give you less latitude to save heat. This would probably be worth measuring and writing an arti…

I think it's likely that all lower bins are artificial, so e.g. underclocking a 2.4 GHz down to 2.0 is probably exactly the same as if you bought the 2.0. But yeah, it would be worth measuring.

Ah, I see. I wrote underclock. It's really undervolting that gets you the big win thermally. Underclocking should be done just as a means of achieving a greater undervolt. I just have these two things in the same mental bin.

Re: An Empirical Analysis of Hardware Failures on a Million Consumer PCs

#79
post #77

Earlier quoted context omitted.

I think it's likely that all lower bins are artificial, so e.g. underclocking a 2.4 GHz down to 2.0 is probably exactly the same as if you bought the 2.0. But yeah, it would be worth measuring.

Ah, I see. I wrote underclock. It's really undervolting that gets you the big win thermally. Underclocking should be done just as a means of achieving a greater undervolt. I just have these two things in the same mental bin.

When you underclock properly (with SpeedStep) it also lowers the voltage... probably to the same voltage that the lower-bin processor would use.

Re: An Empirical Analysis of Hardware Failures on a Million Consumer PCs

#80
post #79

Earlier quoted context omitted.

Ah, I see. I wrote underclock. It's really undervolting that gets you the big win thermally. Underclocking should be done just as a means of achieving a greater undervolt. I just have these two things in the same mental bin.

When you underclock properly (with SpeedStep) it also lowers the voltage... probably to the same voltage that the lower-bin processor would use.

There's not just one voltage here. It's a curve. I suspect that the better the processor turned out, the more favorable your curve turns out to be.
Post reply on HN