Edit: After asking a friend seems most of their research is with Chinese official research institutes and Sensetime. It makes sense now.
A cartel of influential datasets are dominating machine learning research
21–30 of 75 posts
Re: A cartel of influential datasets are dominating machine learning research
#22I think they are highly overestimating how much science advanced pattern matching will provide. Certainly no conceptual understanding will ever come from that.
Most stuff is proven by using advanced pattern matching as "intuition", and then breaking problems down into things we can pattern match and prove.
I'm not sure what conceptual understanding you think isn't pattern matching. They are just rules about patterns and interactions of patterns. These rules were developed through pattern matching.
Re: A cartel of influential datasets are dominating machine learning research
#23Too bad they don't cite the paper "The Benchmark Lottery" (M. Dehghani et al, 2021) ( https://arxiv.org/abs/2107.07002 ) > The world of empirical machine learning (ML) strongly relies on benchmarks in order to determine the relative effectiveness of different algorithms and methods. This paper proposes the notion of "a benchmark lottery" that describes the overall fragility of the ML benchmarking process. The benchma…
Seems like a distinct issue to me, though there are obvious parellels and similar outcomes.
Re: A cartel of influential datasets are dominating machine learning research
#24Somewhat related to this: What methods do people normally use to measure the quality of a dataset? For example, if there are 50 datasets of historical weather data how can I determine which one is garbage?
Re: A cartel of influential datasets are dominating machine learning research
#25Are these big "dominant" institutions charging for this data? No, the spend a lot of resources putting them together and give them away free.
Are they preventing others from giving away data? No, but it costs a lot and they bear that cost.
Are they forcing smaller institutions to use their data? No. Its just a free resource they offer.
Do they get the grants themselves because they have some kind of proprietary access? No, the whole point is that the benchmark is open and everyone has access.
So they collect this data, vet it, propose their use for benchmarks and give it away free. What is the complaint? The problem is not even posed as "well, this data is overfitted in papers or we are solving narrower problems". [Edit: they sort of do, but leaving this here so the comment below makes sense as a follow up.]
No the complaint is that these handful of institutions are giving away free data so too many people use it? Can we have more "problems" of this nature?
Re: A cartel of influential datasets are dominating machine learning research
#26A little confused: Are these big "dominant" institutions charging for this data? No, the spend a lot of resources putting them together and give them away free. Are they preventing others from giving away data? No, but it costs a lot and they bear that cost. Are they forcing smaller institutions to use their data? No. Its just a free resource they offer. Do they get the grants themselves because they have some kind o…
A large institution can make a dataset for X then browbeat other researchers into using X and citing X. Using X also likely leads to citations of derivative work by the lead institution.
Re: A cartel of influential datasets are dominating machine learning research
#27https://paperswithcode.com/dataset/debatesum
https://huggingface.co/datasets/Hellisotherpeople/DebateSum
https://scholar.google.com/citations?user=uHozRV4AAAAJ&hl=en
Re: A cartel of influential datasets are dominating machine learning research
#28A little confused: Are these big "dominant" institutions charging for this data? No, the spend a lot of resources putting them together and give them away free. Are they preventing others from giving away data? No, but it costs a lot and they bear that cost. Are they forcing smaller institutions to use their data? No. Its just a free resource they offer. Do they get the grants themselves because they have some kind o…
I believe the author is suggesting that large institutions are doing this to earn extra citations - the currency of academia. A large institution can make a dataset for X then browbeat other researchers into using X and citing X. Using X also likely leads to citations of derivative work by the lead institution.
Re: A cartel of influential datasets are dominating machine learning research
#29A little confused: Are these big "dominant" institutions charging for this data? No, the spend a lot of resources putting them together and give them away free. Are they preventing others from giving away data? No, but it costs a lot and they bear that cost. Are they forcing smaller institutions to use their data? No. Its just a free resource they offer. Do they get the grants themselves because they have some kind o…
I mean it does get into that:
"They additionally note that blind adherence to this small number of ‘gold’ datasets encourages researchers to achieve results that are overfitted (i.e. that are dataset-specific and not likely to perform anywhere near as well on real-world data, on new academic or original datasets, or even necessarily on different datasets in the ‘gold standard’)."
Re: A cartel of influential datasets are dominating machine learning research
#30Cartel, really?
Right? I thought so too, but this is the first definition from the American heritage dictionary: > A combination of independent business organizations formed to regulate production, pricing, and marketing of goods by the members. And it does seem to apply ¯\_(ツ)_/¯