Live data from Hacker News

JBIG2 Undetectable Data Corruption: Destroying Our Past, One Character at a Time

circuitousroot.com

11–20 of 38 posts

Re: JBIG2 Undetectable Data Corruption: Destroying Our Past, One Character at a Time

#12
post #5

If this is such a serious problem, please demonstrate at least one example in Google Books where a scan shows the wrong letter or number. That really can’t be hard to do, since there are lots of common books in GB, there are even OCRs available so you could even do the error checking automatically for thousands of books. If you can’t show a single case of JBIG data corruption in GOogle Books, you have absolutely no j…

Wouldn’t you need to have an original scan not compressed in JBIG? I’m not sure I understand the suggestion on how to find it in Google Books…

Re: JBIG2 Undetectable Data Corruption: Destroying Our Past, One Character at a Time

#14
post #5

If this is such a serious problem, please demonstrate at least one example in Google Books where a scan shows the wrong letter or number. That really can’t be hard to do, since there are lots of common books in GB, there are even OCRs available so you could even do the error checking automatically for thousands of books. If you can’t show a single case of JBIG data corruption in GOogle Books, you have absolutely no j…

Wouldn’t you need to have an original scan not compressed in JBIG? I’m not sure I understand the suggestion on how to find it in Google Books…

Presumably OCR on the Google Books scans then compare to some known text from a different source (Gutenberg or something)

Re: JBIG2 Undetectable Data Corruption: Destroying Our Past, One Character at a Time

#16

This potential problem is overstated if the quality of the scans is decent. I have scanned literally hundreds of books, and encoded them in both djvu and pdf/jbig2 formats, and have never once found a bad character. (Yes, having heard of this issue, I did initially try to find examples) The free djvu tools have parameters you can adjust to make them more or less aggressive at combining similar characters, and above 2…

The problem lies therein that JBIG2 has a compression factor, that is controlled by people who have an incentive to increase it, where few people know what it actually does. And how much testing is really done when is changed? Will the test include text in other scripts, like Chinese or Japanse? It can go wrong in so many ways.

Jpeg artifacts on text are ugly but they tend to stand out. You will see artifacts long before the text becomes unreadable or ambiguous, and that is a good thing.

Re: JBIG2 Undetectable Data Corruption: Destroying Our Past, One Character at a Time

#17

This potential problem is overstated if the quality of the scans is decent. I have scanned literally hundreds of books, and encoded them in both djvu and pdf/jbig2 formats, and have never once found a bad character. (Yes, having heard of this issue, I did initially try to find examples) The free djvu tools have parameters you can adjust to make them more or less aggressive at combining similar characters, and above 2…

My choice would be the recently standardized JPEG XL designed to replace all of: JPEG, JPEG 2000, PNG, GIF. Among other things, it is supposed to be the first choice for long-term storage. I guess mass adoption will begin after including it in the PDF standard. It is already experimentally in Chrome and Firefox.

Re: JBIG2 Undetectable Data Corruption: Destroying Our Past, One Character at a Time

#19
Sorry, this article is low on content and high on fearmongering. This issue was known years ago and the only example of it seems to be one implementation from Xerox, which was also set to lossy mode --- and I believe that was subsequently fixed too.

The other comments here have linked to the previous articles about this, which do give far more detailed information about the problem. JBIG2 in lossless mode won't do this.

Re: JBIG2 Undetectable Data Corruption: Destroying Our Past, One Character at a Time

#20

Sorry, this article is low on content and high on fearmongering. This issue was known years ago and the only example of it seems to be one implementation from Xerox, which was also set to lossy mode --- and I believe that was subsequently fixed too. The other comments here have linked to the previous articles about this, which do give far more detailed information about the problem. JBIG2 in lossless mode won't do th…

The according talk is linked here in the comments - no, it wasn't just lossy mode. It was much harder, but in the end it was reproduced on other quality levels.

But if it happened once, who's gonna guarantee this won't happen again? When it happens with numbers, worst case it can have fatal consequences.

Post reply on HN