JBIG2 Undetectable Data Corruption: Destroying Our Past, One Character at a Time
21–30 of 38 posts
Re: JBIG2 Undetectable Data Corruption: Destroying Our Past, One Character at a Time
#22This potential problem is overstated if the quality of the scans is decent. I have scanned literally hundreds of books, and encoded them in both djvu and pdf/jbig2 formats, and have never once found a bad character. (Yes, having heard of this issue, I did initially try to find examples) The free djvu tools have parameters you can adjust to make them more or less aggressive at combining similar characters, and above 2…
My choice would be the recently standardized JPEG XL designed to replace all of: JPEG, JPEG 2000, PNG, GIF. Among other things, it is supposed to be the first choice for long-term storage. I guess mass adoption will begin after including it in the PDF standard. It is already experimentally in Chrome and Firefox.
Re: JBIG2 Undetectable Data Corruption: Destroying Our Past, One Character at a Time
#23This potential problem is overstated if the quality of the scans is decent. I have scanned literally hundreds of books, and encoded them in both djvu and pdf/jbig2 formats, and have never once found a bad character. (Yes, having heard of this issue, I did initially try to find examples) The free djvu tools have parameters you can adjust to make them more or less aggressive at combining similar characters, and above 2…
My choice would be the recently standardized JPEG XL designed to replace all of: JPEG, JPEG 2000, PNG, GIF. Among other things, it is supposed to be the first choice for long-term storage. I guess mass adoption will begin after including it in the PDF standard. It is already experimentally in Chrome and Firefox.
Re: JBIG2 Undetectable Data Corruption: Destroying Our Past, One Character at a Time
#24Earlier quoted context omitted.
My choice would be the recently standardized JPEG XL designed to replace all of: JPEG, JPEG 2000, PNG, GIF. Among other things, it is supposed to be the first choice for long-term storage. I guess mass adoption will begin after including it in the PDF standard. It is already experimentally in Chrome and Firefox.
JPEG is optimized for compressing photos, not text or high contrast images.
Re: JBIG2 Undetectable Data Corruption: Destroying Our Past, One Character at a Time
#25I don't belive that's correct; it's rather the other way round.
https://github.com/barak/djvulibre/blob/release.3.5.28/libdj...
> JB2 has strong similarities with the forthcoming JBIG2 standard developed by the "ISO/IEC JTC1 SC29 Working Group 1" which is responsible for both the JPEG and JBIG standards. This is hardly surprising since JB2 was our own proposal for the JBIG2 standard and remained the only proposal for years. The full JBIG2 standard however is significantly more complex and slighlty less efficient than JB2 because it addresses a broader range of applications.
Re: JBIG2 Undetectable Data Corruption: Destroying Our Past, One Character at a Time
#26If this is such a serious problem, please demonstrate at least one example in Google Books where a scan shows the wrong letter or number. That really can’t be hard to do, since there are lots of common books in GB, there are even OCRs available so you could even do the error checking automatically for thousands of books. If you can’t show a single case of JBIG data corruption in GOogle Books, you have absolutely no j…
It is quite common on archive.org. Google know what they are doing and don't have these error.
Re: JBIG2 Undetectable Data Corruption: Destroying Our Past, One Character at a Time
#27Re: JBIG2 Undetectable Data Corruption: Destroying Our Past, One Character at a Time
#28Sorry, this article is low on content and high on fearmongering. This issue was known years ago and the only example of it seems to be one implementation from Xerox, which was also set to lossy mode --- and I believe that was subsequently fixed too. The other comments here have linked to the previous articles about this, which do give far more detailed information about the problem. JBIG2 in lossless mode won't do th…
The according talk is linked here in the comments - no, it wasn't just lossy mode. It was much harder, but in the end it was reproduced on other quality levels. But if it happened once, who's gonna guarantee this won't happen again? When it happens with numbers, worst case it can have fatal consequences.
Re: JBIG2 Undetectable Data Corruption: Destroying Our Past, One Character at a Time
#29Re: JBIG2 Undetectable Data Corruption: Destroying Our Past, One Character at a Time
#30https://media.ccc.de/v/31c3_-_6558_-_de_-_saal_g_-_201412282... Immediately this talk from David Kriesel comes to mind. :)