Live data from Hacker News

The 50 million dollar lie

garyrubinstein.teachforus.org

11–20 of 159 posts

Re: The 50 million dollar lie

#11
post #7

I know nothing about the domain of teacher measurement, and have no opinion on it. But the first chart in the blog post, with the big mass of blue on it, is surely not the strong evidence of weak correlation, that the author is making it out to be? There could still be a strong correlation in that data, even if it looks like a blob of blue - because there are so many data points on that chart, that we can no longer t…

Exactly- It's pointless to counter a misleading graph with another misleading graph.

Re: The 50 million dollar lie

#12
post #7

I know nothing about the domain of teacher measurement, and have no opinion on it. But the first chart in the blog post, with the big mass of blue on it, is surely not the strong evidence of weak correlation, that the author is making it out to be? There could still be a strong correlation in that data, even if it looks like a blob of blue - because there are so many data points on that chart, that we can no longer t…

Correlation means how well the points fit on the line. It is unfortunate that the author did not include a regression line or R^2 value but none the less one can tell there is no line which most of the points lie near. That data is absolutely not strongly correlated.

You seem to be confusing correlation and significance. Having lots of data points makes the correlation more significant, it does not make the correlation stronger.

Re: The 50 million dollar lie

#14
Uh, how does the first graph not have a correlation? As a first order guess, that looks like a very strong correlation to me, probably at least .10 R^2 if you fit a regression to it. I'm guessing the author avoided fitting a regression to the original data because it actually is significant.

Not to mention, it's just plain incorrect statistics if you don't do weighted least squares regression (instead of OLS) on the averaged data set.

Re: The 50 million dollar lie

#15
There are two problems with this article. First is that the "evidence" is from a scatter plot graph, with no statistical measurements attached to it. Stats really aren't something that you can rely on your eyes on. To me, it looks like a weak, positive correlation between previous year's score and current years score. The thing to look for is greater densities in the NE and SW quadrants than the NW and SE. Having a zero-mean "blob" in the data makes it tricky on the eye, but NE and SW points are the ones the model "correctly" predicts, while the NW and SE plots are the ones the models "gets wrong". I think the issue is the eye is looking for a line to process a slope of, but such data doesn't streach into a line. What it should look for is: "given a good score last year, is a good score given this year" or "given a bad score last year, is a bad score given again".

The second issue is that the author does not offer a better solution to the problem. Some information is almost always better than no informantion in decision making. Statistics in human-based samples are always tricky, and its tough to get strong correlations easily. As an aside, this is part of why clinical trials are so tricky and expensive.

(I am not a stats professional, but am a biomedical research scientist with some training in human-population based stats).

Re: The 50 million dollar lie

#16
I don't know enough about statistics to judge the strength of Mr Rubinstein's critique, but I find this a little odd: The axes on his graph is a 'raw score' between -1 and 1, but the axes on the MET study graph is 'standard deviations'. Is that significant?

Re: The 50 million dollar lie

#17

Typically when someone posts a baiting title like that I know not to trust them. Scholarship and evidence don't need hyperbolic headlines.

I think you are mistaken. In today's attention economy even the very best scholarship and the very best evidence-based conclusions will be completely ignored unless there is a hook to get people, preferably lots of people, to read it.

Even here on the allegedly rational Hacker News I've seen item after item sink without trace, whereas others with significantly fewer details and significantly worse evidence get voted up and read by thousands.

Alex Bellos recently passed on to me what his editor told him. It doesn't matter how good your article is, if no one reads it, or if no one gets to the point, it's no use and a waste of your time.

I'd like to think "Scholarship and evidence don't need hyperbolic headlines." Sadly, I think that scholarship and evidence say otherwise.

Re: The 50 million dollar lie

#18

This is a surprising blog post in that I draw the complete opposite conclusion the author does. The author seems to think that the averaging hides volatility (which it does), which leads to incorrect conclusions drawn. Whereas to me it looks like it removes visual noise to show an actual trend. In his original scatter plot, because of the big ball in the middle it's easy to handwave and say, "look a random blob" -- b…

The other thing the author does is to derive absolute findings when the author himself admits to making an assumption about how the numbers were obtained. To me, using the word "lie" in the title is therefore nothing more than click-baiting. And on a separate note - there's a lot of anti-Gates sentiment flying around lately. Whether his methods are the best or not, he's using his own money . I find it really difficul…

Bill Gates's position as a famous wealthy guy means he has more of a responsibility than the rest of us to not put out misleading graphs. Even if he is spreading misinformation with his own money.

Re: The 50 million dollar lie

#19
post #7

I know nothing about the domain of teacher measurement, and have no opinion on it. But the first chart in the blog post, with the big mass of blue on it, is surely not the strong evidence of weak correlation, that the author is making it out to be? There could still be a strong correlation in that data, even if it looks like a blob of blue - because there are so many data points on that chart, that we can no longer t…

Correlation means how well the points fit on the line. It is unfortunate that the author did not include a regression line or R^2 value but none the less one can tell there is no line which most of the points lie near. That data is absolutely not strongly correlated. You seem to be confusing correlation and significance. Having lots of data points makes the correlation more significant, it does not make the correlati…

one can tell there is no line which most of the points lie near

No, that's not correct. You can't conclude anything from looking at a big blob because you don't know the density of points at different places in the blob. This is the point the guy you replied to was making.

As an extreme example, imagine a billion data points that fit perfectly on a straight line. Then superimpose a million data points randomly on top of it. What does it look like? A big blob. But almost every point is highly correlated with that straight line.

Re: The 50 million dollar lie

#20
He doesn't actually say what the positive correlation is, which I find disingenuous and suspicious. Because his scatterplot is so non-detailed, it's about as useful as the following scatterplot of SAT scores:

X|X

---

X|X

Where the axes are scores above and below 600 in math on the first and the second time taking the test. There are some individuals who do better or worse - but just because we can draw a detail-obfuscating graph of data doesn't mean there is no detail in the data, or no correlation, or that SAT scores won't help us predict future SAT scores. It just means we can draw a detail-obfuscating graph, and that there's not a perfect correlation.

One point the author does have is that these scores are not perfectly predictive, so bad performance one year shouldn't mean that the teacher gets i.e. fired. OK. He seems to be protesting using information to draw any conclusions at all. Perhaps it would be more effective to show what fallacious or irrational conclusions are being falsely drawn from this data and used to fire teachers, and then object to those specific instances of irrational actions. This value-added measure has not been shown to be intrinsically unreliable or wrongheaded but maybe some applications of it are.

Post reply on HN