Live data from Hacker News

Reproducible machine learning with PyTorch and Quilt

blog.paperspace.com

11–20 of 27 posts

Re: Reproducible machine learning with PyTorch and Quilt

#12
post #3

Oh, this resonates with me so much! I'm running 4 different DeepSpeech models right now, each using a differently processed version of LibriSpeech dataset (mfcc/fbanks/linear spectrograms, deltas? energy? padding? etc). Because the original DS papers didn't bother describing it, and every implementation I found uses completely different methods and libraries. Not to mention every one of those implementation packages…

Why don't you use STFT + Conv2D like Deep Speech 2 did. It works well in my case.

Re: Reproducible machine learning with PyTorch and Quilt

#14
post #5

Was not aware of Quilt for hosting datasets. Is it the go-to in this area? What other alternatives are there?

I've tried dat (https://datproject.org/) and git lfs (https://git-lfs.github.com/) but so far have found quilt to be easiest to use & best fitting to my use case (experimental physics characterization experiments).

Re: Reproducible machine learning with PyTorch and Quilt

#15
post #5

Was not aware of Quilt for hosting datasets. Is it the go-to in this area? What other alternatives are there?

I have been researching data orchestration/versioning tools for a long time and have been following the Quilt guys closely. It is definitely one of the more powerful tools in the ML/AI engineer's toolbox and solves a huge problem that almost everyone runs in to right out the gate. It's still early days in this space but Quilt gets a lot of things right and I'm super excited to see this product develop.

Full disclosure: I run Paperspace (https://www.paperspace.com) and am working with the Quilt team to integrate their tools in to our platform.

Re: Reproducible machine learning with PyTorch and Quilt

#17
post #5

Was not aware of Quilt for hosting datasets. Is it the go-to in this area? What other alternatives are there?

Just use AWS S3 (or similar) and shell scripts. My team uses a git repository named something like "data-packages", which is nothing but a collection of shell scripts with the name .sh, that perform the necessary download and extraction steps to get a dataset from S3. Data sets are immutable by convention, so any changes to a data set requires you to provide a totally new shell script. That script could download an older data set and then mutate it if you don't want to maintain large copies of big data sets, but the older data set itself is not permitted to be mutated on S3.

My team has found this drastically easier than Quilt, and we do a ton of stuff with reproducible environments in Docker, creating Makefiles to reproduce exact model training with the exact same data, etc. We probably hit just about every case there is (huge models, small models, models where we'd like to train separately or collectively on a bunch of different benchmark data sets, in-house data sets, models that need to be refreshed with new data in pipelines, etc.) So far, Quilt has not been competitive with a simple repo of shell scripts for us, in terms of ease of use or effectiveness in maintaining different packages of data.

The other super nice thing is that when people start out on new models or experiments, we already have our in-house maintained copies of a bunch of academic data sets, private data sets, etc., and you can throw together an incredibly simple Dockerfile or Makefile that uses the appropriate script. It's just one or two lines of shell code and voila, you have an environment with the dataset you want. Check that into git and now your experiment is immediately reproducible from day one. We've found this to dramatically increase the amount of code review that researchers engage in for checking their statistical methodology and sanity checking their intended models or experiments. With Quilt, you have the extra issue of versioning (rather than harshly enforcing all data sets to be immutable ... even just adding one more training example to the data set means you must provide a new shell script that downloads the old data, injects your lone additional sample, and has a documentation entry about exactly what it is doing), as well as the overhead of using yet another tool instead of super standard shell scripts.

For me, any of the tools that pop up attempting to be like conda-forge but for data packages is sort of like taking a gatling gun to a problem that can be solved with a hammer.

Re: Reproducible machine learning with PyTorch and Quilt

#18
post #5

Was not aware of Quilt for hosting datasets. Is it the go-to in this area? What other alternatives are there?

Just use AWS S3 (or similar) and shell scripts. My team uses a git repository named something like "data-packages", which is nothing but a collection of shell scripts with the name .sh, that perform the necessary download and extraction steps to get a dataset from S3. Data sets are immutable by convention, so any changes to a data set requires you to provide a totally new shell script. That script could download an o…

Do you store the datasets as tar/zip archives on S3, or do you have some way of representing how a collection of items goes together to form a dataset?

Re: Reproducible machine learning with PyTorch and Quilt

#19
post #5

Was not aware of Quilt for hosting datasets. Is it the go-to in this area? What other alternatives are there?

Just use AWS S3 (or similar) and shell scripts. My team uses a git repository named something like "data-packages", which is nothing but a collection of shell scripts with the name .sh, that perform the necessary download and extraction steps to get a dataset from S3. Data sets are immutable by convention, so any changes to a data set requires you to provide a totally new shell script. That script could download an o…

Interesting thoughts. Quilt has a ways to grow. You correctly point out that, in some cases, S3 is lighter weight. You'll see future versions of Quilt get lighter, and offer more S3-like "just store this" functionality. In its next minor revision, Quilt simplifies point updates (i.e. it will be possible to update a single training example without materializing the entire package).

That said, there are a few areas where your system glosses over the needs of a data pipeline:

* "immutable by convention" is not a data preservation strategy; the system should enforce immutability

* what about deserialization? it's not enough to store and move bits. there are so many examples of "serdes" headaches. pickling (yes, pickle is a horrible format) in python 2 vs python 3 is one example. not to mention performance. my point is not that scripts can't do serdes, but that serdes information should travel with the data, so it's (mostly) transparent to the consumer.

* multiple writers (e.g. suppose you are generating training data in a distributed manner) requires write atomicity at the bucket level, which S3 doesn't provide

* deduplication of data fragments - I can see how one might do this with a "scripts over S3" strategy, but it's complicated enough that it's far easier to rely on a third-party app that just works in this regard

* fine-grained permissions - what if each data package has a different audience? sure, you can roll this with S3, but is that the best use of developer time?

* change history and access auditing

* querying and filtering - in many cases there is an enormous data corpus which needs to be sliced a different way by each user, e.g. Google Open Images. it is much more robust to have a single query mechanism that understands data layout than to write a fresh script for each slice.

* indexing data so they are searchable, etc.

PS - I am a contributor to Quilt.

Re: Reproducible machine learning with PyTorch and Quilt

#20
post #18

Earlier quoted context omitted.

Just use AWS S3 (or similar) and shell scripts. My team uses a git repository named something like "data-packages", which is nothing but a collection of shell scripts with the name .sh, that perform the necessary download and extraction steps to get a dataset from S3. Data sets are immutable by convention, so any changes to a data set requires you to provide a totally new shell script. That script could download an o…

Do you store the datasets as tar/zip archives on S3, or do you have some way of representing how a collection of items goes together to form a dataset?

Quilt co-founder here, no, Quilt doesn’t use tar or zip. Each package version has a manifest that specifies the set of items it contains. Each item in the collection is stored and transported separately. Items are identified by their hash and stored once even if used in more than one package version.
Post reply on HN