Live data from Hacker News

Peeking Inside Gigantic Zips with Only Kilobytes

ritiksahni.com

1–10 of 30 posts

Re: Peeking Inside Gigantic Zips with Only Kilobytes

#3
For implementation in a library, you can use HttpRangeReader [1][2] in zip.js [3] (disclaimer: I am the author). It's a solid feature that has been in the library for about 10 years.

[1] https://gildas-lormeau.github.io/zip.js/api/classes/HttpRang...

[2] https://github.com/gildas-lormeau/zip.js/blob/master/tests/a...

[3] https://github.com/gildas-lormeau/zip.js

Re: Peeking Inside Gigantic Zips with Only Kilobytes

#4
This is really cool! Could also make a useful standalone command line tool.

I think the general pattern - using the range header + prior knowledge of a file format to only download the parts of a file that are relevant - is still really underutilized.

One small problem I see is that a server that does not support range requests would just try to send you the entire file in the first request, I think.

So maybe doing a preflight HEAD request first to see if the server sends back Accept-Ranges could be useful.

https://developer.mozilla.org/en-US/docs/Web/HTTP/Guides/Ran...

Re: Peeking Inside Gigantic Zips with Only Kilobytes

#8
post #3

For implementation in a library, you can use HttpRangeReader [1][2] in zip.js [3] (disclaimer: I am the author). It's a solid feature that has been in the library for about 10 years. [1] https://gildas-lormeau.github.io/zip.js/api/classes/HttpRang... [2] https://github.com/gildas-lormeau/zip.js/blob/master/tests/a... [3] https://github.com/gildas-lormeau/zip.js

Based on your experience, is zip the optimal archive format for long term digital archival in object storage if the use case calls for reading archives via http for scanning and cherry picking? Or is there a more optimal archive format?

Re: Peeking Inside Gigantic Zips with Only Kilobytes

#9
> That question took me into the guts of the ZIP format, where I learned there’s a tiny index at the end that points to everything else.

Tangential, but any Free Software that uses `shared-mime-info` to identify files (any of your GNOMEs, KDEs, etc) are unable to correctly identify Zip files by their EOCD due to lack of accepted syntax for defining search patterns based on negative file offsets. Please show your support on this Issue if you would also like to see this resolved: https://gitlab.freedesktop.org/xdg/shared-mime-info/-/issues... (linking to my own comment, so no this is not brigading)

Anything using `file(1)` does not have this problem: https://github.com/file/file/blob/280e121/magic/Magdir/zip#L...

Re: Peeking Inside Gigantic Zips with Only Kilobytes

#10
post #4

This is really cool! Could also make a useful standalone command line tool. I think the general pattern - using the range header + prior knowledge of a file format to only download the parts of a file that are relevant - is still really underutilized. One small problem I see is that a server that does not support range requests would just try to send you the entire file in the first request, I think. So maybe doing a…

How common is it in practice today to not support ranges? I remember back in the early days of broadband (c. 2000) when having a Download Manager was something most nerds endorsed, that most servers then supported partial downloads. Aside from toy projects has anyone encountered a server which didn't allow ranges (unless specifically configured to forbid it)?
Post reply on HN