Live data from Hacker News

Pigz: Parallel gzip for modern multi-processor, multi-core machines

zlib.net

111–120 of 197 posts

Re: Pigz: Parallel gzip for modern multi-processor, multi-core machines

#111

Earlier quoted context omitted.

This episode was fascinating. I had heard of LZ4 but not Zstd. It spurred me to make changes to our system at work that are reducing file sizes by as much as 25%. It’s great to have a podcast in which I learn practical stuff!

Probably the most underrated feature of zstd (likely because it's so unusual) is the ability to create a separate compression dictionary. This allows you to develop customized and highly efficient dictionaries that are highly specific to a type of data AND allow you to compress elements of that data without including an entire separate dictionary in every compression output. So for example take logfiles. You can trai…

I saw that zstd and brotli both suppport creating custom dictionaries but I couldn't find any tutorials showing how to do this. Perhaps you could share code?

Re: Pigz: Parallel gzip for modern multi-processor, multi-core machines

#112
post #76

I heard of pigz in the discussions following my interview of Yann Collet, creator of LZ4 and zstd. If you'll excuse the plug, here is the LZ4 story: Yann was bored and working as a project manager. So he started working on a game for his old HP 48 graphing calculator. Eventually, this hobby led him to revolutionize the field of data compression, releasing LZ4, ZStandard, and Finite State Entropy coders. His code ende…

Fastest open source compression algorithms. RAD game tools have proprietary ones that are faster and have better compression ratios, but since you have to pay for a license, they will never be widespread.

I've been following RAD for a long time and I love Charles Bloom's blog. They are proprietary but he also makes a lot of code public. For instance, he showed how RAD switched from arithmetic coders to Assymetrical Number Systems and added code.

Re: Pigz: Parallel gzip for modern multi-processor, multi-core machines

#113

Earlier quoted context omitted.

Probably the most underrated feature of zstd (likely because it's so unusual) is the ability to create a separate compression dictionary. This allows you to develop customized and highly efficient dictionaries that are highly specific to a type of data AND allow you to compress elements of that data without including an entire separate dictionary in every compression output. So for example take logfiles. You can trai…

I do this to make extraordinarily small UDP packets for a low latency system. I record the raw payload then build a dictionary for the data, then share it on both sides. It reduces the packet overhead by removing the dictionary and it does a much better job than other approaches.

I saw that zstd and brotli both suppport creating custom dictionaries but I couldn't find any tutorials showing how to do this. Perhaps you could share code?

Re: Pigz: Parallel gzip for modern multi-processor, multi-core machines

#114
post #106
post #53

Unless the recipient of whatever you are compressing absolutely requires gzip, you should not use gzip or pigz. Instead you should use zstd as it compresses faster, decompresses faster, and yields smaller files. It also supports parallelism (via “-T”) which supplants the pigz use case. There literally are no trade-offs; it is better in every objective way. In 2023, friends don’t let friends use gzip.

Not really. "The recipients" are for example millions of browsers that don't understand zstd.

I agree. I wanted to use Brotli for my startup because it allows creating custom dictionaries but I had to resort to gzip because Brotli was difficult to setup on my CDN.

Re: Pigz: Parallel gzip for modern multi-processor, multi-core machines

#115

I heard of pigz in the discussions following my interview of Yann Collet, creator of LZ4 and zstd. If you'll excuse the plug, here is the LZ4 story: Yann was bored and working as a project manager. So he started working on a game for his old HP 48 graphing calculator. Eventually, this hobby led him to revolutionize the field of data compression, releasing LZ4, ZStandard, and Finite State Entropy coders. His code ende…

One wild thing is how much performance wins were available compared to ZLib. Pigz is parrellel, but what if you just had a better way to compress and decompress than DEFLATE? When zstd came out – and Brotli before it to a certain extent – they were 3x faster than ZLib with a slightly higher compression ratio. You'd think that such performance jumps in something as well explored as data compression would be hard to co…

zlib is above all portable, and runs in small memory footprints. Size vs space and platform specific functionality all have costs associated with them.

Re: Pigz: Parallel gzip for modern multi-processor, multi-core machines

#116
post #33

Earlier quoted context omitted.

You'd be surprised. There are some workloads - for me, it's geospatial data - where bzip2 clobbers all of the alternatives.

I'm using bzip2 to compress a specific type of backup. In my case I cannot afford to steal CPU time from other running processes, so the backup process runs with severely limited CPU percentage. By crude testing I found that bzip2 used 10x less memory and finished several times faster than xz, while being very close on the compression rate. Other algorithms like zstd and gz resulted in much lower compression rates. I…

I've not seen one that picks the best compression algorithm, but I've seen ones that perform a test to try and determine if it is worth compressing. For example borg-backup software can be configured to try a light & fast compression algorithm on each chunk of data. If the chunk is compressed, then it uses a more computationally expensive algorithm to really squash it down.

Re: Pigz: Parallel gzip for modern multi-processor, multi-core machines

#117

Earlier quoted context omitted.

Probably the most underrated feature of zstd (likely because it's so unusual) is the ability to create a separate compression dictionary. This allows you to develop customized and highly efficient dictionaries that are highly specific to a type of data AND allow you to compress elements of that data without including an entire separate dictionary in every compression output. So for example take logfiles. You can trai…

I saw that zstd and brotli both suppport creating custom dictionaries but I couldn't find any tutorials showing how to do this. Perhaps you could share code?

Basically,

`zstd --train `

will output a dictionary file, and then the `-D ` option when used for either compression or decompression will then use that dictionary first.

You can also investigate "man zstd" or google "zstd --train" for more details. The directory for the training must consist of many small files each of which is an example artifact; if you want to split, say, a single log file into files of each line, you can use, say, a bash script like this (note that I just created this with ChatGPT and eyeballed it, it looks correct but I haven't run it yet!): https://gist.github.com/pmarreck/91124e761e45d6860834eb046d6... (Also, don't forget to set it as executable with `chmod +x split_file.bash` before you try to run it directly)

Re: Pigz: Parallel gzip for modern multi-processor, multi-core machines

#118

John Carmack just had a tweet today on this problem: >I started a tar with bzip command on a big directory, and it has been running for two days. Of course, it is only using 1.07 cores out of the 128 available. The Unix pipeline tool philosophy often isn’t aligned with parallel performance. https://twitter.com/ID_AA_Carmack/status/1656708636570271768...

Shame he didn't discover pbzip2 before starting that job.

If it's less than 98% complete he could still stop it, start over and still finish sooner.

Re: Pigz: Parallel gzip for modern multi-processor, multi-core machines

#119
post #42

I'm a big fan of pigz. I use it in my home-grown backup script for my Linux laptop. It can compress the incremental tar output from my filesystem snapshot fast enough to saturate the I/O to my external USB3 hard drive. This is a low bar, but single-threaded gzip (or bzip2) could not do it!

Give xz and zstd a go. You'll love them.

Re: Pigz: Parallel gzip for modern multi-processor, multi-core machines

#120
post #66
post #53

Unless the recipient of whatever you are compressing absolutely requires gzip, you should not use gzip or pigz. Instead you should use zstd as it compresses faster, decompresses faster, and yields smaller files. It also supports parallelism (via “-T”) which supplants the pigz use case. There literally are no trade-offs; it is better in every objective way. In 2023, friends don’t let friends use gzip.

gzip has the advantage of being ubiquitous. It's pretty much guaranteed to be available on every modern Unix-alike. And is good enough for most purposes. Zstd is getting there but I personally don't bother with it on a daily basis except in situations where both performance and compression ratio are important, like build artifact pipelines or large archives.

You make it sound like installing zstd is a big deal. Which it is not.
Post reply on HN