Earlier quoted context omitted.
Why do you think it works this way and not the other way around as well? Grocery stores shop their data around to see who will pay the most for it. A person who is out of work goes to the local courthouse and requests a bunch of records, compiles them into a spreadsheet and then cold (or warm with something like LinkedIn) calls to see if anyone is interested in the data. An online quiz company is going out of busines…
Good question. I thought it works that way because it takes a lot of work to sanitize and cross-link people's data to other datasets accurately, so even if it's a "push" model, I still can't believe that every single website that does this does their own data cleaning & ML & whatnot. It's far too much repeated work and a good business to just do the work and sell it off to others. So I'd assume a few companies have t…
It's far too much repeated work
Companies will repeat work over and over again if it's cheaper than buying it, they have custom needs that aren't filled with the data available, etc. Businesses repeat work all the time, and this is not any different. Additionally, for many businesses in the sector, they themselves are the primary source for data. For them it's not repeated work.
a good business to just do the work and sell it off to others.
Yes, that's why some aggregators exist. They make money by brokering the data from multiple sources, some primary and some resold. But they are the tip of the iceberg.
You seem to be under the impression that there is some small list of companies who are all working from primary sources, and that everyone then gets feeds of data from those companies. This would make sense if gathering data was very difficult, or had a natural resource-like limitations. So that model works well for something like diamond mining (as compared to diamond growing), because the number of diamond mines are limited, and there is a natural entry barrier. However, that doesn't take into account the fact that gathering this data is generally easy. Sometimes it's very easy, such as a sftp feed of data from a government records database. Sometimes it's a bit harder, such as needing to physically be present to obtain the data.
That means there is very little barrier to entry, and thus generally there is going to be a lot of competition, and thus many companies vying to make money.
Personal data has value just like any other commodity. So a bit of economic theory goes a long way to understanding what the boundaries of a market might be. Low production cost, high profit goods generally have a large number of companies in the market.