Peeking Inside Gigantic Zips with Only Kilobytes
11–20 of 30 posts
Re: Peeking Inside Gigantic Zips with Only Kilobytes
#12This is really cool! Could also make a useful standalone command line tool. I think the general pattern - using the range header + prior knowledge of a file format to only download the parts of a file that are relevant - is still really underutilized. One small problem I see is that a server that does not support range requests would just try to send you the entire file in the first request, I think. So maybe doing a…
How common is it in practice today to not support ranges? I remember back in the early days of broadband (c. 2000) when having a Download Manager was something most nerds endorsed, that most servers then supported partial downloads. Aside from toy projects has anyone encountered a server which didn't allow ranges (unless specifically configured to forbid it)?
For static files served by CDNs or an "established" HTTP servers I think support is pretty much a given (though e.g. Python's FastAPI only got support in 2020 [1]), but for anything dynamic, I doubt many devs would go through the trouble and implement support if it wasn't strictly necessary for their usecase.
E.g. the URL may point to a service endpoint that loads the file contents from a database or blob storage instead of the file system. Then the service would have to implement range support itself and translate them to the necessary storage/database calls (if those exist), etc etc. That's some effort you have to put in.
Even for static files, there may be reverse proxies in front that (unintentionally) remove the support again. E.g. [2]
[1] https://github.com/Kludex/starlette/issues/950
[2] https://caddy.community/t/cannot-seek-further-in-videos-usin...
Re: Peeking Inside Gigantic Zips with Only Kilobytes
#137-zip does this. You can see it if you open (to view) a large ZIP file on slow network drive. There's no way it is downloading the whole thing. You can extract single files from the ZIP also with only a little traffic.
Re: Peeking Inside Gigantic Zips with Only Kilobytes
#14For implementation in a library, you can use HttpRangeReader [1][2] in zip.js [3] (disclaimer: I am the author). It's a solid feature that has been in the library for about 10 years. [1] https://gildas-lormeau.github.io/zip.js/api/classes/HttpRang... [2] https://github.com/gildas-lormeau/zip.js/blob/master/tests/a... [3] https://github.com/gildas-lormeau/zip.js
Based on your experience, is zip the optimal archive format for long term digital archival in object storage if the use case calls for reading archives via http for scanning and cherry picking? Or is there a more optimal archive format?
Re: Peeking Inside Gigantic Zips with Only Kilobytes
#15Re: Peeking Inside Gigantic Zips with Only Kilobytes
#16This is also quite easy to do with .tar files, not to be confused with .tar.gz files.
Re: Peeking Inside Gigantic Zips with Only Kilobytes
#17For implementation in a library, you can use HttpRangeReader [1][2] in zip.js [3] (disclaimer: I am the author). It's a solid feature that has been in the library for about 10 years. [1] https://gildas-lormeau.github.io/zip.js/api/classes/HttpRang... [2] https://github.com/gildas-lormeau/zip.js/blob/master/tests/a... [3] https://github.com/gildas-lormeau/zip.js
Based on your experience, is zip the optimal archive format for long term digital archival in object storage if the use case calls for reading archives via http for scanning and cherry picking? Or is there a more optimal archive format?
1) The format has limited and archaic support for file metadata - e.g. file modification times are stored as a MS-DOS timestamp with a 2-second (!) resolution, and there's no standard system for representing other metadata.
2) The single-level central directory can be awkward to work with for archives containing a very large number of members.
3) Support for 64-bit file sizes exists but is a messy hack.
4) Compression operates on each file as a separate stream, reducing its effectiveness for archives containing many small files. The format does support pluggable compression methods, but there's no straightforward way to support "solid" compression.
5) There is technically no way to reliably identify a ZIP file, as the end of central directory record can appear at any location near the end of the file, and the file can contain arbitrary data at its start. Most tools recognize ZIP files by the presence of a local file header at the start ("PK\x01\x02"), but that's not reliable.
Re: Peeking Inside Gigantic Zips with Only Kilobytes
#18Earlier quoted context omitted.
Based on your experience, is zip the optimal archive format for long term digital archival in object storage if the use case calls for reading archives via http for scanning and cherry picking? Or is there a more optimal archive format?
ZIP isn't a terrible format, but it has a couple of flaws and limitations which make it a less than ideal format for long-term archiving. The biggest ones I'd call out are: 1) The format has limited and archaic support for file metadata - e.g. file modification times are stored as a MS-DOS timestamp with a 2-second (!) resolution, and there's no standard system for representing other metadata. 2) The single-level cen…
I do it by ignoring ZIP's native compression entirely, using store-only ZIP files and then compressing the whole thing at the filesystem level instead.
Here's an example comparison of the same WWW site rip in a DEFLATE ZIP, in a store-only ZIP with zstd filesystem compression, in a tar with same zstd filesystem compression (identical size but less useful for seeking due to lack of trailing directory versus ZIP), and finally the raw size pre-zipping:
982M preserve.mactech.com.deflate.zip
408M preserve.mactech.com.store.zip
410M preserve.mactech.com.tar
3.8G preserve.mactech.com
[Lammy@popola] zfs get compression spinthedisc/Backups/WWW
NAME PROPERTY VALUE SOURCE
spinthedisc/Backups/WWW compression zstd local
This probably wouldn't help GP with their need for HTTP seeking since their HTTP server would incur a decompress+recompress at the filesystem boundary.Re: Peeking Inside Gigantic Zips with Only Kilobytes
#19Re: Peeking Inside Gigantic Zips with Only Kilobytes
#20I'll dig up a link.