Earlier quoted context omitted.
> data ownership and pipelines not ML. By your argument, the money is in data hoarding and brokering, and renting that data (with DRM) to ML outfits, not doing ML. Anyway, cleaning dirty data isn't execeptionally hard, it's just boring work; getting the raw data is the hard part.
This is an interesting observation that deserves to be addressed. I think in practice, what happens is that there's no easy transactional boundary that can keep the data ownership and ML in separate firms. The theory of the firm [1] states that firms arise when transaction costs are less than the economic inefficiencies of centralized resource allocation. There are some pretty heavy transaction costs between doing da…
There is. Data broking is a well established business model in finance. Every bank and hedge fund is running ML models on data licensed from a provider such as Reuter’s, S&P, etc there are dozens of such brokers.