Live data from Hacker News

Déjà vu: Fast efficient probabilistic deduplication

github.com

1–10 of 13 posts

Re: Déjà vu: Fast efficient probabilistic deduplication

#3
post #2

Poor name choice. https://en.wikipedia.org/wiki/D%C3%A9j%C3%A0_Vu_(software) https://github.com/worldveil/dejavu https://github.com/appbaseio/dejavu https://github.com/IndigoUnited/js-dejavu

Yeah, should have put more thought into it then I did. But I guess its to late now.

Re: Déjà vu: Fast efficient probabilistic deduplication

#4
post #2

Poor name choice. https://en.wikipedia.org/wiki/D%C3%A9j%C3%A0_Vu_(software) https://github.com/worldveil/dejavu https://github.com/appbaseio/dejavu https://github.com/IndigoUnited/js-dejavu

I must say his is the most-fitting use of the name, though.

Re: Déjà vu: Fast efficient probabilistic deduplication

#7
post #2

Poor name choice. https://en.wikipedia.org/wiki/D%C3%A9j%C3%A0_Vu_(software) https://github.com/worldveil/dejavu https://github.com/appbaseio/dejavu https://github.com/IndigoUnited/js-dejavu

You missed one! One of my favourite adventure games!

https://en.wikipedia.org/wiki/D%C3%A9j%C3%A0_Vu_(video_game)

Re: Déjà vu: Fast efficient probabilistic deduplication

#8
Bloom filters don't seem useful as the primary means of deduplication for an actual data storage system to me.

- False positives means marking data as duplicate when it's not.

- Bloom filters are not associative. So unless the application is very special and can find/identify/retrieve the data later in some other (likely inefficient?) way, a separate index is required anyway. But if you have a separate index of the data you already have, you can just use that to deduplicate. This index is the main memory cost of deduplicating storage.

Bloom filters are of course interesting in certain scenarios e.g. to statistically reduce network traffic.

Re: Déjà vu: Fast efficient probabilistic deduplication

#10

Could you use this to deduplicate files in a filesystem? Could you use this as the deduplication method in ZFS?

> Could you use this to deduplicate files in a filesystem?

Depends on what you mean and your constraints. Deduplicate file entries if you can live with a rare false positives, sure. A setup of 1million entrie limit with 1/1billion false positive chance will get ~80M mem usage.

If you want to check for duplicate files, probably not.

> Could you use this as the deduplication method in ZFS?

I would say no.

Post reply on HN