Was not aware of Quilt for hosting datasets. Is it the go-to in this area? What other alternatives are there?
Reproducible machine learning with PyTorch and Quilt
11–20 of 27 posts
Re: Reproducible machine learning with PyTorch and Quilt
#12Oh, this resonates with me so much! I'm running 4 different DeepSpeech models right now, each using a differently processed version of LibriSpeech dataset (mfcc/fbanks/linear spectrograms, deltas? energy? padding? etc). Because the original DS papers didn't bother describing it, and every implementation I found uses completely different methods and libraries. Not to mention every one of those implementation packages…
Re: Reproducible machine learning with PyTorch and Quilt
#13Was not aware of Quilt for hosting datasets. Is it the go-to in this area? What other alternatives are there?
Re: Reproducible machine learning with PyTorch and Quilt
#14Was not aware of Quilt for hosting datasets. Is it the go-to in this area? What other alternatives are there?
Re: Reproducible machine learning with PyTorch and Quilt
#15Was not aware of Quilt for hosting datasets. Is it the go-to in this area? What other alternatives are there?
Full disclosure: I run Paperspace (https://www.paperspace.com) and am working with the Quilt team to integrate their tools in to our platform.
Re: Reproducible machine learning with PyTorch and Quilt
#16Inference example: https://www.paperspace.com/console/jobs/js4mqzm91fj2lg
Disclosure: I work on Paperspace
Re: Reproducible machine learning with PyTorch and Quilt
#17Was not aware of Quilt for hosting datasets. Is it the go-to in this area? What other alternatives are there?
My team has found this drastically easier than Quilt, and we do a ton of stuff with reproducible environments in Docker, creating Makefiles to reproduce exact model training with the exact same data, etc. We probably hit just about every case there is (huge models, small models, models where we'd like to train separately or collectively on a bunch of different benchmark data sets, in-house data sets, models that need to be refreshed with new data in pipelines, etc.) So far, Quilt has not been competitive with a simple repo of shell scripts for us, in terms of ease of use or effectiveness in maintaining different packages of data.
The other super nice thing is that when people start out on new models or experiments, we already have our in-house maintained copies of a bunch of academic data sets, private data sets, etc., and you can throw together an incredibly simple Dockerfile or Makefile that uses the appropriate script. It's just one or two lines of shell code and voila, you have an environment with the dataset you want. Check that into git and now your experiment is immediately reproducible from day one. We've found this to dramatically increase the amount of code review that researchers engage in for checking their statistical methodology and sanity checking their intended models or experiments. With Quilt, you have the extra issue of versioning (rather than harshly enforcing all data sets to be immutable ... even just adding one more training example to the data set means you must provide a new shell script that downloads the old data, injects your lone additional sample, and has a documentation entry about exactly what it is doing), as well as the overhead of using yet another tool instead of super standard shell scripts.
For me, any of the tools that pop up attempting to be like conda-forge but for data packages is sort of like taking a gatling gun to a problem that can be solved with a hammer.
Re: Reproducible machine learning with PyTorch and Quilt
#18Was not aware of Quilt for hosting datasets. Is it the go-to in this area? What other alternatives are there?
Just use AWS S3 (or similar) and shell scripts. My team uses a git repository named something like "data-packages", which is nothing but a collection of shell scripts with the name .sh, that perform the necessary download and extraction steps to get a dataset from S3. Data sets are immutable by convention, so any changes to a data set requires you to provide a totally new shell script. That script could download an o…
Re: Reproducible machine learning with PyTorch and Quilt
#19Was not aware of Quilt for hosting datasets. Is it the go-to in this area? What other alternatives are there?
Just use AWS S3 (or similar) and shell scripts. My team uses a git repository named something like "data-packages", which is nothing but a collection of shell scripts with the name .sh, that perform the necessary download and extraction steps to get a dataset from S3. Data sets are immutable by convention, so any changes to a data set requires you to provide a totally new shell script. That script could download an o…
That said, there are a few areas where your system glosses over the needs of a data pipeline:
* "immutable by convention" is not a data preservation strategy; the system should enforce immutability
* what about deserialization? it's not enough to store and move bits. there are so many examples of "serdes" headaches. pickling (yes, pickle is a horrible format) in python 2 vs python 3 is one example. not to mention performance. my point is not that scripts can't do serdes, but that serdes information should travel with the data, so it's (mostly) transparent to the consumer.
* multiple writers (e.g. suppose you are generating training data in a distributed manner) requires write atomicity at the bucket level, which S3 doesn't provide
* deduplication of data fragments - I can see how one might do this with a "scripts over S3" strategy, but it's complicated enough that it's far easier to rely on a third-party app that just works in this regard
* fine-grained permissions - what if each data package has a different audience? sure, you can roll this with S3, but is that the best use of developer time?
* change history and access auditing
* querying and filtering - in many cases there is an enormous data corpus which needs to be sliced a different way by each user, e.g. Google Open Images. it is much more robust to have a single query mechanism that understands data layout than to write a fresh script for each slice.
* indexing data so they are searchable, etc.
PS - I am a contributor to Quilt.
Re: Reproducible machine learning with PyTorch and Quilt
#20Earlier quoted context omitted.
Just use AWS S3 (or similar) and shell scripts. My team uses a git repository named something like "data-packages", which is nothing but a collection of shell scripts with the name .sh, that perform the necessary download and extraction steps to get a dataset from S3. Data sets are immutable by convention, so any changes to a data set requires you to provide a totally new shell script. That script could download an o…
Do you store the datasets as tar/zip archives on S3, or do you have some way of representing how a collection of items goes together to form a dataset?