Live data from Hacker News

JBIG2 Undetectable Data Corruption: Destroying Our Past, One Character at a Time

circuitousroot.com

1–10 of 38 posts

Re: JBIG2 Undetectable Data Corruption: Destroying Our Past, One Character at a Time

#2
This potential problem is overstated if the quality of the scans is decent. I have scanned literally hundreds of books, and encoded them in both djvu and pdf/jbig2 formats, and have never once found a bad character. (Yes, having heard of this issue, I did initially try to find examples) The free djvu tools have parameters you can adjust to make them more or less aggressive at combining similar characters, and above 200DPI it's hard to even force the issue to happen for something like an 8 versus a 6. That was my experience anyway.

It's also telling that the example trotted out is always a decades-old photocopier implementation. Unless someone does some comparisons of recent scans and finds a bunch of altered characters, I just don't buy this concern. MRC+JBig2 is really great and the size savings have not been trivial for me personally.

(I hesitate to add that the alternative would be, I assume, DCT/jpeg page images, which introduce lots of noisy artifacts of their own... destroying our history one DCT-block at a time!)

Re: JBIG2 Undetectable Data Corruption: Destroying Our Past, One Character at a Time

#3

This potential problem is overstated if the quality of the scans is decent. I have scanned literally hundreds of books, and encoded them in both djvu and pdf/jbig2 formats, and have never once found a bad character. (Yes, having heard of this issue, I did initially try to find examples) The free djvu tools have parameters you can adjust to make them more or less aggressive at combining similar characters, and above 2…

There are a number of scans I've come across on archive.org where the person uploading has used jbig2 and there are pages where letters like 'e' get swapped for an 's'. My general rule when creating pdfs with scans is: If it's mostly text or line drawings and it will be used in a professional setting use png/flate. If it's a photo use jpg (especially if the source is a jpg).

Re: JBIG2 Undetectable Data Corruption: Destroying Our Past, One Character at a Time

#5
If this is such a serious problem, please demonstrate at least one example in Google Books where a scan shows the wrong letter or number.

That really can’t be hard to do, since there are lots of common books in GB, there are even OCRs available so you could even do the error checking automatically for thousands of books.

If you can’t show a single case of JBIG data corruption in GOogle Books, you have absolutely no justification in calling it such a serious problem!

Re: JBIG2 Undetectable Data Corruption: Destroying Our Past, One Character at a Time

#6

This potential problem is overstated if the quality of the scans is decent. I have scanned literally hundreds of books, and encoded them in both djvu and pdf/jbig2 formats, and have never once found a bad character. (Yes, having heard of this issue, I did initially try to find examples) The free djvu tools have parameters you can adjust to make them more or less aggressive at combining similar characters, and above 2…

No post body was provided.

Re: JBIG2 Undetectable Data Corruption: Destroying Our Past, One Character at a Time

#7

This potential problem is overstated if the quality of the scans is decent. I have scanned literally hundreds of books, and encoded them in both djvu and pdf/jbig2 formats, and have never once found a bad character. (Yes, having heard of this issue, I did initially try to find examples) The free djvu tools have parameters you can adjust to make them more or less aggressive at combining similar characters, and above 2…

There are a number of scans I've come across on archive.org where the person uploading has used jbig2 and there are pages where letters like 'e' get swapped for an 's'. My general rule when creating pdfs with scans is: If it's mostly text or line drawings and it will be used in a professional setting use png/flate. If it's a photo use jpg (especially if the source is a jpg).

I'd love to see some examples (especially if the scan appeared to be of good quality and not a grainy 100dpi affair). Do you have any idea what kinds of documents you were browsing when you found them?

Re: JBIG2 Undetectable Data Corruption: Destroying Our Past, One Character at a Time

#8

This potential problem is overstated if the quality of the scans is decent. I have scanned literally hundreds of books, and encoded them in both djvu and pdf/jbig2 formats, and have never once found a bad character. (Yes, having heard of this issue, I did initially try to find examples) The free djvu tools have parameters you can adjust to make them more or less aggressive at combining similar characters, and above 2…

There are a number of scans I've come across on archive.org where the person uploading has used jbig2 and there are pages where letters like 'e' get swapped for an 's'. My general rule when creating pdfs with scans is: If it's mostly text or line drawings and it will be used in a professional setting use png/flate. If it's a photo use jpg (especially if the source is a jpg).

There are both lossy and lossless jbig2 variants.

Re: JBIG2 Undetectable Data Corruption: Destroying Our Past, One Character at a Time

#9
post #8

Earlier quoted context omitted.

There are a number of scans I've come across on archive.org where the person uploading has used jbig2 and there are pages where letters like 'e' get swapped for an 's'. My general rule when creating pdfs with scans is: If it's mostly text or line drawings and it will be used in a professional setting use png/flate. If it's a photo use jpg (especially if the source is a jpg).

There are both lossy and lossless jbig2 variants.

Unless something has changed, Google books has always used the lossy encoding. Quoting the original paper:

    The most obvious way to compress pages is to losslessly encode them as a single 1 bpp image (see Figure 2). However, we get much better compression by using symbol encoding and accepting some loss of image data. Although our compression is lossy it is not clear how much information is actually being lost - the letter forms on a page are obviously supposed to be uniform in shape, the variation comes from printing errors and lack of resolution in image capture.
Probably only a small fraction of errors are JBIG2 substitutions though.

Re: JBIG2 Undetectable Data Corruption: Destroying Our Past, One Character at a Time

#10
post #5

If this is such a serious problem, please demonstrate at least one example in Google Books where a scan shows the wrong letter or number. That really can’t be hard to do, since there are lots of common books in GB, there are even OCRs available so you could even do the error checking automatically for thousands of books. If you can’t show a single case of JBIG data corruption in GOogle Books, you have absolutely no j…

It is quite common on archive.org. Google know what they are doing and don't have these error.
Post reply on HN