Live data from Hacker News

Reproducible machine learning with PyTorch and Quilt

blog.paperspace.com

21–27 of 27 posts

Re: Reproducible machine learning with PyTorch and Quilt

#21
post #19

Earlier quoted context omitted.

Just use AWS S3 (or similar) and shell scripts. My team uses a git repository named something like "data-packages", which is nothing but a collection of shell scripts with the name .sh, that perform the necessary download and extraction steps to get a dataset from S3. Data sets are immutable by convention, so any changes to a data set requires you to provide a totally new shell script. That script could download an o…

Interesting thoughts. Quilt has a ways to grow. You correctly point out that, in some cases, S3 is lighter weight. You'll see future versions of Quilt get lighter, and offer more S3-like "just store this" functionality. In its next minor revision, Quilt simplifies point updates (i.e. it will be possible to update a single training example without materializing the entire package). That said, there are a few areas whe…

> “immutable by convention" is not a data preservation strategy; the system should enforce immutability”

I actually disagree with this. In a Python-like “consenting adults” philosophy, I think it’s worse to spend engineering effort to guarantee immutability rather than to trust people not to and just have a reasonable system of backups.

Immutability by convention is 99.99999999% as good as enforced immutability for this particular type of task, and there’s even less risk with a good backup strategy to fall back if there is an accident.

Change history and access auditing are super easy on S3, as is fine-grained access control. With immutability by convention, change history is just git history, and you can customize access groups on a file-by-file basis if you want. You could also instrument logging in the shell scripts themselves if you really want, and I’m not convinced that’s worse than a third party doing it, especially if your logging backend is prone to change, or you wantbyo pipe stuff to Grafana, etc., which are quite common needs.

Querying and filtering are separate post-processing tasks. They should be expressed as source code that mutates a data set after downloading a local copy, and in fact version controlling any data cleaning, post-processing, etc., should be kept completely separate from the management of a data package. They are logically hugely different parts of the process. Slicing data ought to be up to the individual developer or researcher, to choose their tools, to optimize, etc. Version control of that source code is the right way to make that part of the work reproducible, not trying to tie custom treatments into a version of a data package.

Your points about dedup and deserialization are good ones. I can imagine problem cases for a simple script approach, but I can also say even for gigantic in-house image data sets, creating multiple slightly different materialized copies has rarely been an issue.

Re: Reproducible machine learning with PyTorch and Quilt

#22
post #18

Earlier quoted context omitted.

Just use AWS S3 (or similar) and shell scripts. My team uses a git repository named something like "data-packages", which is nothing but a collection of shell scripts with the name .sh, that perform the necessary download and extraction steps to get a dataset from S3. Data sets are immutable by convention, so any changes to a data set requires you to provide a totally new shell script. That script could download an o…

Do you store the datasets as tar/zip archives on S3, or do you have some way of representing how a collection of items goes together to form a dataset?

For some data sets we use an archive format and the corresponding script unpackages the data. For others, we have all individual files in S3. Using archives improves download speed, and sometimes we provide both, so some narrowly-defined data sets benefit from the archive file, while others can pick and choose specific samples to include or exclude.

One of the nice things about the simple script approach is that the dataset can be defined however you like. Whatever the script retrieves and unpacks for you, that is the data set of that script. It could be a superset or subset of other data, and intersection of samples with a certain property from other data. As long as that definition is treated as immutable, and the backing data is immutable, it defines a specific collection of items however you want.

Re: Reproducible machine learning with PyTorch and Quilt

#23
post #3

Oh, this resonates with me so much! I'm running 4 different DeepSpeech models right now, each using a differently processed version of LibriSpeech dataset (mfcc/fbanks/linear spectrograms, deltas? energy? padding? etc). Because the original DS papers didn't bother describing it, and every implementation I found uses completely different methods and libraries. Not to mention every one of those implementation packages…

Why don't you use STFT + Conv2D like Deep Speech 2 did. It works well in my case.

The DeepSpeech2 paper does not include any details about audio processing. I see an older Baidu-Research implementation of DS1 that uses "log of linear spectrogram from FFT energy". Also, there's a pytorch implementation [1], where they use Librosa's STFT, is that what you're referring to?

That's two more implementations that I haven't considered. I'm sure most of the processing steps under the hood are the same or similar, but as I'm not an audio processing expert, I can't tell which method is better (and why).

And it's hard to tell if it "works well" because or despite the way I processed the files.

[1] https://github.com/SeanNaren/deepspeech.pytorch

Re: Reproducible machine learning with PyTorch and Quilt

#25

A step in the right direction for machine learning in science, but they could have done some more research into naming conflicts: $ apt-cache show quilt Package: quilt [..] Description-en: Tool to work with series of patches Quilt manages a series of patches by keeping track of the changes each of them makes. They are logically organized as a stack, and you can apply, un-apply, refresh them easily by traveling into t…

i hear you. on pypi the name is uncontested so, at least in the python eco-system, there is only one quilt. that said, for future revisions we'll try for a unique name because it can indeed be confusing, e.g. in the apt-get case.

Re: Reproducible machine learning with PyTorch and Quilt

#27
post #5

Was not aware of Quilt for hosting datasets. Is it the go-to in this area? What other alternatives are there?

Has anyone tried out the option of self-hosting Quilt registries? I really like the idea of Quilt, although I am worried that my network bandwidth would be an issue for 10-100GB datasets...
Post reply on HN