DataChain: DBT for Unstructured Data
github.com
DataChain: DBT for Unstructured Data
1–10 of 27 posts
Re: DataChain: DBT for Unstructured Data
#2How does one wrangle terabytes of data on a local machine?
Re: DataChain: DBT for Unstructured Data
#3> It is made to organize your unstructured data into datasets and wrangle it at scale on your local machine. How does one wrangle terabytes of data on a local machine?
(that's is different from let's say DVC - that does copy files into a local cache, always)
Re: DataChain: DBT for Unstructured Data
#4It doesn't really replace any of the tooling we use to wrangle data at scale (like prefect or dagster or temporal) but as a local library it seems to be excellent, I think what confused me most was the comparison to dbt.
I like the from_* utils and the magic of the Column class operator overloading and how chains can be used as datasets. Love how easy checkpointing is too. Will give it a go
Re: DataChain: DBT for Unstructured Data
#5Maintainer and author here. Happy to answer any questions.
We built DataChain because our DVC couldn't fully handle data transformations and versioning directly in S3/GCS/Azure without data copying.
Analogy with "DBT for unstractured data" applies very well to DataChain since it transforms data (using Python, not SQL) inside in storages (S3, not DB). Happy to talk more!
Re: DataChain: DBT for Unstructured Data
#6It took me a minute to grok what this was for, but I think I like it It doesn't really replace any of the tooling we use to wrangle data at scale (like prefect or dagster or temporal) but as a local library it seems to be excellent, I think what confused me most was the comparison to dbt. I like the from_* utils and the magic of the Column class operator overloading and how chains can be used as datasets. Love how ea…
Try it out - looking forward to your feedback!
Re: DataChain: DBT for Unstructured Data
#7My most common use cases involve getting PDFs or HTML files and I have to parse the metadata to store along with the embedding.
Would I have to run a process to extract file metadata into JSONs for every embedding/chunk? Would keys created based off document be title+chunk_no?
Very interested in this because documents from clients are subject to random changes and I don’t have very robust systems in place.
Re: DataChain: DBT for Unstructured Data
#8Cool! Does this assume the unstructured data already has a corresponding metadata file? My most common use cases involve getting PDFs or HTML files and I have to parse the metadata to store along with the embedding. Would I have to run a process to extract file metadata into JSONs for every embedding/chunk? Would keys created based off document be title+chunk_no? Very interested in this because documents from clients…
Extract metadata as usual, then return the result as JSON or a Pydantic object. DataChian will automatically serialize it to internal dataset structure (SQLite), which can be exported to CSV/Parquet.
In case of PDF/HTML, you will likely produce multiple documents per file which is also supported - just `yield return my_result` multiple times from map().
Check out video: https://www.youtube.com/watch?v=yjzcPCSYKEo Blog post: https://datachain.ai/blog/datachain-unstructured-pdf-process...
Re: DataChain: DBT for Unstructured Data
#9> It is made to organize your unstructured data into datasets and wrangle it at scale on your local machine. How does one wrangle terabytes of data on a local machine?
The idea is that it doesn't store binary files locally, just pointers in the DB + meta data (SQLite if you run locally, open source). So, it's versioning, structuring of datasets, etc by "references" if you wish. (that's is different from let's say DVC - that does copy files into a local cache, always)
Re: DataChain: DBT for Unstructured Data
#10Earlier quoted context omitted.
The idea is that it doesn't store binary files locally, just pointers in the DB + meta data (SQLite if you run locally, open source). So, it's versioning, structuring of datasets, etc by "references" if you wish. (that's is different from let's say DVC - that does copy files into a local cache, always)
So in the case from the README, where you're trying to curate a sample of your data, the only thing that you're reading is the metadata, UNTIL you run `export_files` and that actually copies the binary data to your local machine?
This way, you might end up downloading just 1% of your data, as defined by the metadata filter.