Live data from Hacker News

A cartel of influential datasets are dominating machine learning research

unite.ai

21–30 of 75 posts

Re: A cartel of influential datasets are dominating machine learning research

#21
I'm extremely surprised to see CUHK on the list. I'm living 10 minutes from them and never knew they are a big player in the ML scene. A brief search online didn't show up anything interesting apart from the CelebA dataset.

Edit: After asking a friend seems most of their research is with Chinese official research institutes and Sensetime. It makes sense now.

Re: A cartel of influential datasets are dominating machine learning research

#22
post #7

I think they are highly overestimating how much science advanced pattern matching will provide. Certainly no conceptual understanding will ever come from that.

I think you are highly underestimating how much science and math comes mainly from advanced pattern matching.

Most stuff is proven by using advanced pattern matching as "intuition", and then breaking problems down into things we can pattern match and prove.

I'm not sure what conceptual understanding you think isn't pattern matching. They are just rules about patterns and interactions of patterns. These rules were developed through pattern matching.

Re: A cartel of influential datasets are dominating machine learning research

#23

Too bad they don't cite the paper "The Benchmark Lottery" (M. Dehghani et al, 2021) ( https://arxiv.org/abs/2107.07002 ) > The world of empirical machine learning (ML) strongly relies on benchmarks in order to determine the relative effectiveness of different algorithms and methods. This paper proposes the notion of "a benchmark lottery" that describes the overall fragility of the ML benchmarking process. The benchma…

Seems like a distinct issue to me, though there are obvious parellels and similar outcomes.

Agree, I was not saying it was redundant. But I would have expected a reference, as that paper was the first related item that came to my mind.

Re: A cartel of influential datasets are dominating machine learning research

#24
post #4

Somewhat related to this: What methods do people normally use to measure the quality of a dataset? For example, if there are 50 datasets of historical weather data how can I determine which one is garbage?

I would say it's garbage if there's a paper like this: https://arxiv.org/abs/1902.01007

Re: A cartel of influential datasets are dominating machine learning research

#25
A little confused:

Are these big "dominant" institutions charging for this data? No, the spend a lot of resources putting them together and give them away free.

Are they preventing others from giving away data? No, but it costs a lot and they bear that cost.

Are they forcing smaller institutions to use their data? No. Its just a free resource they offer.

Do they get the grants themselves because they have some kind of proprietary access? No, the whole point is that the benchmark is open and everyone has access.

So they collect this data, vet it, propose their use for benchmarks and give it away free. What is the complaint? The problem is not even posed as "well, this data is overfitted in papers or we are solving narrower problems". [Edit: they sort of do, but leaving this here so the comment below makes sense as a follow up.]

No the complaint is that these handful of institutions are giving away free data so too many people use it? Can we have more "problems" of this nature?

Re: A cartel of influential datasets are dominating machine learning research

#26

A little confused: Are these big "dominant" institutions charging for this data? No, the spend a lot of resources putting them together and give them away free. Are they preventing others from giving away data? No, but it costs a lot and they bear that cost. Are they forcing smaller institutions to use their data? No. Its just a free resource they offer. Do they get the grants themselves because they have some kind o…

I believe the author is suggesting that large institutions are doing this to earn extra citations - the currency of academia.

A large institution can make a dataset for X then browbeat other researchers into using X and citing X. Using X also likely leads to citations of derivative work by the lead institution.

Re: A cartel of influential datasets are dominating machine learning research

#27
Well, I recently released a NLP dataset of almost 200K documents and I got 4 whole citations in a year. Wish I could find a way to join this cartel and get someone to use it or even care

https://paperswithcode.com/dataset/debatesum

https://huggingface.co/datasets/Hellisotherpeople/DebateSum

https://scholar.google.com/citations?user=uHozRV4AAAAJ&hl=en

Re: A cartel of influential datasets are dominating machine learning research

#28
post #26

A little confused: Are these big "dominant" institutions charging for this data? No, the spend a lot of resources putting them together and give them away free. Are they preventing others from giving away data? No, but it costs a lot and they bear that cost. Are they forcing smaller institutions to use their data? No. Its just a free resource they offer. Do they get the grants themselves because they have some kind o…

I believe the author is suggesting that large institutions are doing this to earn extra citations - the currency of academia. A large institution can make a dataset for X then browbeat other researchers into using X and citing X. Using X also likely leads to citations of derivative work by the lead institution.

Is the term "browbeat" fair here? I don't think they're making calls and saying "Oh, nice paper, but I notice you used this other dataset..." No, they're putting out good quality datasets that people want to use. If that earns them a citation, good for them. Its the least I can do for helping me test my algo.

Re: A cartel of influential datasets are dominating machine learning research

#29

A little confused: Are these big "dominant" institutions charging for this data? No, the spend a lot of resources putting them together and give them away free. Are they preventing others from giving away data? No, but it costs a lot and they bear that cost. Are they forcing smaller institutions to use their data? No. Its just a free resource they offer. Do they get the grants themselves because they have some kind o…

>The problem is not even posed as "well, this data is overfitted in papers or we are solving narrower problems".

I mean it does get into that:

"They additionally note that blind adherence to this small number of ‘gold’ datasets encourages researchers to achieve results that are overfitted (i.e. that are dataset-specific and not likely to perform anywhere near as well on real-world data, on new academic or original datasets, or even necessarily on different datasets in the ‘gold standard’)."

Re: A cartel of influential datasets are dominating machine learning research

#30
post #2

Cartel, really?

Right? I thought so too, but this is the first definition from the American heritage dictionary: > A combination of independent business organizations formed to regulate production, pricing, and marketing of goods by the members. And it does seem to apply ¯\_(ツ)_/¯

They're not regulating anything. They just make the best datasets and those are the ones that get used
Post reply on HN