"A new research report demonstrates the connection between the rising amount of mass produced housing featuring a dedicated sex dungeon and biased content in the training set used for the compression algorithm in a popular line of architectural plotters"
JBIG2 Undetectable Data Corruption: Destroying Our Past, One Character at a Time
31–38 of 38 posts
Re: JBIG2 Undetectable Data Corruption: Destroying Our Past, One Character at a Time
#32Sorry, this article is low on content and high on fearmongering. This issue was known years ago and the only example of it seems to be one implementation from Xerox, which was also set to lossy mode --- and I believe that was subsequently fixed too. The other comments here have linked to the previous articles about this, which do give far more detailed information about the problem. JBIG2 in lossless mode won't do th…
The problem is the existence and use of a codec for a purpose, that harms that purpose. It doesn't matter what it's name is.
Are scans being produced with this codec? And is it not true that not only is there data loss or corruption like with jpg, but that unlike jpg artifacts you can sometimes not know that the data loss or corruption occurred?
That does make it not merely a problem of "I wish they scanned this in higher quality" but a far worse problem of "this document looks good so I trust it" when in fact it was corrupt.
Re: JBIG2 Undetectable Data Corruption: Destroying Our Past, One Character at a Time
#33Earlier quoted context omitted.
There are a number of scans I've come across on archive.org where the person uploading has used jbig2 and there are pages where letters like 'e' get swapped for an 's'. My general rule when creating pdfs with scans is: If it's mostly text or line drawings and it will be used in a professional setting use png/flate. If it's a photo use jpg (especially if the source is a jpg).
I'd love to see some examples (especially if the scan appeared to be of good quality and not a grainy 100dpi affair). Do you have any idea what kinds of documents you were browsing when you found them?
Re: JBIG2 Undetectable Data Corruption: Destroying Our Past, One Character at a Time
#34This potential problem is overstated if the quality of the scans is decent. I have scanned literally hundreds of books, and encoded them in both djvu and pdf/jbig2 formats, and have never once found a bad character. (Yes, having heard of this issue, I did initially try to find examples) The free djvu tools have parameters you can adjust to make them more or less aggressive at combining similar characters, and above 2…
Re: JBIG2 Undetectable Data Corruption: Destroying Our Past, One Character at a Time
#35Earlier quoted context omitted.
There are a number of scans I've come across on archive.org where the person uploading has used jbig2 and there are pages where letters like 'e' get swapped for an 's'. My general rule when creating pdfs with scans is: If it's mostly text or line drawings and it will be used in a professional setting use png/flate. If it's a photo use jpg (especially if the source is a jpg).
I'd love to see some examples (especially if the scan appeared to be of good quality and not a grainy 100dpi affair). Do you have any idea what kinds of documents you were browsing when you found them?
There's apparently 22 pages of fluff at the beggining, so you should add that the page numbers. Also "the third line" was meant to be the 3rd line from the bottom.
Direct links (to be opened in DjView):
https://ia800204.us.archive.org/28/items/cu31924026442156/cu...
https://ia800204.us.archive.org/28/items/cu31924026442156/cu...
Re: JBIG2 Undetectable Data Corruption: Destroying Our Past, One Character at a Time
#36Most probably its raison d'être is for non-critical stuff, such as literature, where character flips may be easily detectable during encoding and after the fact (eg. with a spell-checker) and won't possibly subvert the meaning of a whole literary work anyway.
Re: JBIG2 Undetectable Data Corruption: Destroying Our Past, One Character at a Time
#37Wait till you get ML based compression and you get replacements with entirely different but plausibly looking alternatives that go beyond single character substitutions. "A new research report demonstrates the connection between the rising amount of mass produced housing featuring a dedicated sex dungeon and biased content in the training set used for the compression algorithm in a popular line of architectural plott…
Re: JBIG2 Undetectable Data Corruption: Destroying Our Past, One Character at a Time
#38Earlier quoted context omitted.
I'd love to see some examples (especially if the scan appeared to be of good quality and not a grainy 100dpi affair). Do you have any idea what kinds of documents you were browsing when you found them?
https://news.ycombinator.com/item?id=17435514 (found via https://en.wikipedia.org/wiki/DjVu ) There's apparently 22 pages of fluff at the beggining, so you should add that the page numbers. Also "the third line" was meant to be the 3rd line from the bottom. Direct links (to be opened in DjView): https://ia800204.us.archive.org/28/items/cu31924026442156/cu... https://ia800204.us.archive.org/28/items/cu31924026442156/c…
So, I have to conclude that the original conversions were done with very aggressive settings (or very bad jbig2 implementations). If that's they way the internet archive did all its encodings, then I admit there must be several problem documents, and they should re-do the conversions from the original JP2k files they seem to have kept. Not sure if google books made similar mistakes.