> Tons and tons of companies are trying to solve this. You'll get lost among all the vendors trying to capitalize on the “Modern Data Stack”. They're expensive.
Yes, and Yes. However I see a path for a new kind of warehouse that isn't a "Git for Data" or a Databricks' product.
"Topics" (Kafka terminology) containing versioned datasets could help organise and discover what is available in a given repo. The Data versioning would help ensuring reproducibility of any AI pipeline consuming those, because artefacts would be kept unaltered "forever". And finally, availability would be supported by a simple set of read-only replicas.
Your best alternative today consist of stuffing a blob storage with parquet files, and giving write access to a proxy (or to everyone, if you want to live dangerously). Where are append-only semantics ? Conflict-free artefact creation with concurrent reads ?
There is massive opportunities here, it's almost exciting !
PS : > I'm particularly interested in ending up at a database or analytics company. So if you're a database or analytics company hiring managers or developers, feel free to message me!
I see you in the thread, so does a consulting company with a massive Data Science arm is of any interest to you ? :)