Live data from Hacker News

Xerox scanners and photocopiers randomly alter numbers in scanned documents

dkriesel.com

41–50 of 118 posts

Re: Xerox scanners and photocopiers randomly alter numbers in scanned documents

#41
post #34

Ouch, imagine this happens in a hospital with a prescription or something. It could really have some serious implications.

Indeed, I keep a copy of my lab results for the last N years because they sometimes get lost, once through no real fault of the doctor ( http://en.wikipedia.org/wiki/2011_Joplin_tornado ). Grrr, I'm now going to have to view every lab report that's not an original with suspicion, and make sure my doctors aren't making recommendations due to screwed up copies. Lossy compression is not an acceptable default for a gener…

This isn't even lossy compression - it's misleading compression

Re: Xerox scanners and photocopiers randomly alter numbers in scanned documents

#42
post #23

Earlier quoted context omitted.

It is amazing the things people expect to be emailed. Can you email me the MRI scan? It's over 1500 images. Can you email it?

It's amazing to me the engineers who refuse to update their worldviews about normal people's mental models for "sending data" that still get amazed by this. The size of the files or the number of them are totally irrelevant.

Normal people seem to get that it's considerably harder to ship a barn than a letter, and that if you want to move a barn you use a specialty service rather than the post office.

The size and number of files are and should be totally relevant even to "normal" people. When someone asks for something in e-mail, it's perfectly reasonable to say "no, it's much too big" and expect them to understand.

Re: Xerox scanners and photocopiers randomly alter numbers in scanned documents

#43
This should be on the computer risks digest.

There is virtually no reason whatsoever for this problem to exist. This is the domain of "making a problem more risky and complicated than it needs to be" and royally screwing people in the process.

Might as well throw the paperwork in a bin and set fire to it.

Re: Xerox scanners and photocopiers randomly alter numbers in scanned documents

#44
post #36
post #3

I can't quite see the reason why you would lossily compress something when your machine's purpose is to duplicate things. Anyone got a reasonable reason for doing this?

Others have pointed out a credible explanation: to have the document take less space on their hard disk. However, it does not have to be compression, per se. Modern copiers want to correct all kinds of errors such as creases and staples. They also want to optimize the colors. To do that, they have logic for detecting what areas of the page are full-color and which are black and white, which are half-tone printed, whi…

Well we have 14TiB of financial documents archived on our kit. There is no way we even would consider such compression!!!

The whole thing is dangerous and wholly illogical.

This is akin to a crappy crime flick where someone hits the "enhance!" button on a CCTV still a few times and gets to see the dirt on the guy's teeth.

In this case, the computer decides the guy is female and has no teeth.

Re: Xerox scanners and photocopiers randomly alter numbers in scanned documents

#47
post #14

This class of error is called (by me, at least) a "contoot" because, long ago, when I was writing the JBIG2 compressor for Google Books PDFs, the first example was on the contents page of book. The title, "Contents", was set in very heavy type which happened to be an unexpected edge case in the classifier and it matched the "o" with the "e" and "n" and output "Contoots". The classifier was adjusted and these errors m…

[deleted]

Re: Xerox scanners and photocopiers randomly alter numbers in scanned documents

#48
post #47
post #14

This class of error is called (by me, at least) a "contoot" because, long ago, when I was writing the JBIG2 compressor for Google Books PDFs, the first example was on the contents page of book. The title, "Contents", was set in very heavy type which happened to be an unexpected edge case in the classifier and it matched the "o" with the "e" and "n" and output "Contoots". The classifier was adjusted and these errors m…

[deleted]

[deleted]

Re: Xerox scanners and photocopiers randomly alter numbers in scanned documents

#50
post #36

Earlier quoted context omitted.

Others have pointed out a credible explanation: to have the document take less space on their hard disk. However, it does not have to be compression, per se. Modern copiers want to correct all kinds of errors such as creases and staples. They also want to optimize the colors. To do that, they have logic for detecting what areas of the page are full-color and which are black and white, which are half-tone printed, whi…

Well we have 14TiB of financial documents archived on our kit. There is no way we even would consider such compression!!! The whole thing is dangerous and wholly illogical. This is akin to a crappy crime flick where someone hits the "enhance!" button on a CCTV still a few times and gets to see the dirt on the guy's teeth. In this case, the computer decides the guy is female and has no teeth.

[deleted]
Post reply on HN