Live data from Hacker News

Twenty-nine teams use same dataset, find contradicting results [pdf]

osf.io

11–20 of 36 posts

Re: Twenty-nine teams use same dataset, find contradicting results [pdf]

#11
Reminds me of the idea (Robin Hanson's, I think?) to add an extra layer of blindness to studies: during peer review, take the original data, and write a separate paper with the opposite conclusion. Randomize which reviewers get which version. Your original paper is then only accepted if they reject the inverted version.

Re: Twenty-nine teams use same dataset, find contradicting results [pdf]

#12
post #4

Lies, damned lies, and statistics https://en.wikipedia.org/wiki/Lies,_damned_lies,_and_statist... Statistics can be manipulated surprisingly easily.

There are three kinds of lies. There are also three kinds of comments I see in this thread: > "This is interesting, here's some thoughts and ideas that further contribute to this subject" > "This is interesting, here's a link to some further writing on this subject" > "The entire concept and discipline of statistics is bullshit." Par for the course here at Hacker News.

Hey, the null hypothesis is powerful and valuable. I, for one, and happy that all three types are well-represented; all three are healthy in moderation.

I also think that the quote fits in quite nicely here, it's not a wholesale rejection of statistics.

Re: Twenty-nine teams use same dataset, find contradicting results [pdf]

#14
I understand how tempting it is in our age of big data and all that stuff to perceive this as some curious new phenomena, but it really is not. This is precisely the reason why we've come up with some criteria for "science" quite a while ago. And in fact, all this experiment is pretty meaningless.

So, for starters: 29 students get the same question on the math/physics/chemistry exam and give 29 different answers. Breaking news? Obviously not. Either the question was outrageously bad worded (not such a rare thing, sadly), or students didn't do very well and we've got at most 1 correct answer.

Basically, we've got the very same situation here. Except our "students" were doing statistics, which is not really math and not really natural science. Which is why it is somehow "acceptable" to end up with the results like that.

If we are doing math, whatever result we get must be backed up with formally correct proof. Which doesn't mean of course, that 2 good students cannot get contradicting results, but at least one of their proofs is faulty, which can be shown. And this is how we decide what's "correct".

If we are doing science (e.g. physics) our question must be formulated in a such way that it is verifiable by setting up an experiment. If experiment didn't get us what we expected — our theory is wrong. If it did — it might be correct.

Here, our original question was "if players with dark skin tone are more likely than light skin toned players to receive red cards from referees", which is shit, and not a scientific hypothesis. We can define "more likely" as we want. What we really want to know: if during next N matches happening in what we can consider "the same environment" black athletes are going to get more red cards than white athletes. Which is quite obviously a bad idea for a study, because the number of trials we need is too big for so loosely defined setting: not even 1 game will actually happen in isolated environment, players will be different, referees will be different, each game will change the "state" of our world. Somebody might even say that the whole culture has changed since we started the experiment, so obviously whatever the first dataset was — it's no longer relevant.

Statistics is only a tool, not a "science", as some people might (incorrectly) assume. It is not the fault of methods we apply that we get something like that, but rather the discipline that we apply them to. And "results" like that is why physics is accepted as a science, and sociology never really was.

Re: Twenty-nine teams use same dataset, find contradicting results [pdf]

#15
post #14

I understand how tempting it is in our age of big data and all that stuff to perceive this as some curious new phenomena, but it really is not. This is precisely the reason why we've come up with some criteria for "science" quite a while ago. And in fact, all this experiment is pretty meaningless. So, for starters: 29 students get the same question on the math/physics/chemistry exam and give 29 different answers. Bre…

Physics uses statiatics all the time, e.g. detecting the higgs boson at cern. Do you have a formal proof thay each time they fired the accelerator it was going to be i.i.d.?

Re: Twenty-nine teams use same dataset, find contradicting results [pdf]

#16
post #14

I understand how tempting it is in our age of big data and all that stuff to perceive this as some curious new phenomena, but it really is not. This is precisely the reason why we've come up with some criteria for "science" quite a while ago. And in fact, all this experiment is pretty meaningless. So, for starters: 29 students get the same question on the math/physics/chemistry exam and give 29 different answers. Bre…

[deleted]

Re: Twenty-nine teams use same dataset, find contradicting results [pdf]

#17
post #15
post #14

I understand how tempting it is in our age of big data and all that stuff to perceive this as some curious new phenomena, but it really is not. This is precisely the reason why we've come up with some criteria for "science" quite a while ago. And in fact, all this experiment is pretty meaningless. So, for starters: 29 students get the same question on the math/physics/chemistry exam and give 29 different answers. Bre…

Physics uses statiatics all the time, e.g. detecting the higgs boson at cern. Do you have a formal proof thay each time they fired the accelerator it was going to be i.i.d.?

Please, read comments you are answering to.

Re: Twenty-nine teams use same dataset, find contradicting results [pdf]

#18
post #17
post #15

Earlier quoted context omitted.

Physics uses statiatics all the time, e.g. detecting the higgs boson at cern. Do you have a formal proof thay each time they fired the accelerator it was going to be i.i.d.?

Please, read comments you are answering to.

I did.

Re: Twenty-nine teams use same dataset, find contradicting results [pdf]

#19
post #8
post #3

Earlier quoted context omitted.

The blog post is a great overview as well as useful context, thanks for sharing it. TL:DR summary: Scientific results are highly contingent on subjective decisions at the analysis stage. Different (well-founded) data analysis techniques on a fairly simple and well-defined problem can give radically different results. It's very interesting research -- a great real-life example supporting the models Scott Page et. al.…

Off-topic, but you seem well positioned to answer: Why do you say "TL:DR" here when summarizing a short blog post that you enjoyed? Clearly the meaning has diverged from the original abbreviated insult of "Too long; didn't read", but I don't understand what people mean when they use it today. Why did you phrase it this way? Are you a native English speaker? If not intended to be derogatory, does the dissonance bother…

How do you see it as derogatory? I'm a native English speaker and have never thought of it that way. I didn't click the link, but did appreciate his short summary – and upvoted him for it. ;)

Re: Twenty-nine teams use same dataset, find contradicting results [pdf]

#20
post #14

I understand how tempting it is in our age of big data and all that stuff to perceive this as some curious new phenomena, but it really is not. This is precisely the reason why we've come up with some criteria for "science" quite a while ago. And in fact, all this experiment is pretty meaningless. So, for starters: 29 students get the same question on the math/physics/chemistry exam and give 29 different answers. Bre…

Your rant makes no sense. I flip a coin 100 times and it comes up tails 99 times. You are basically saying that asking "Is the coin more likely to come up tails" isn't a real scientific question. That's just silly.
Post reply on HN