Even if Bill Gates isn't doing the best job, every time I see one of these articles, it appears that there isn't any advice on better ways to grade teachers. Did I miss that in the article? Are there a lot of other people in the field that are doing better work? (I would really like to read if anyone can provide links to papers to other ways of analyzing teachers)
Did I miss that in the article The blog link to this thread is a propaganda site. It conspicuously doesn't reference any of the original materials that it is attempting (weakly) to refute. There was an article on HN last week that referenced a more source material article about the project. Here's the link to the METProject's web site: http://www.metproject.org/
The 50 million dollar lie
141–150 of 159 posts
Re: The 50 million dollar lie
#142Earlier quoted context omitted.
Correlation means how well the points fit on the line. It is unfortunate that the author did not include a regression line or R^2 value but none the less one can tell there is no line which most of the points lie near. That data is absolutely not strongly correlated. You seem to be confusing correlation and significance. Having lots of data points makes the correlation more significant, it does not make the correlati…
one can tell there is no line which most of the points lie near No, that's not correct. You can't conclude anything from looking at a big blob because you don't know the density of points at different places in the blob. This is the point the guy you replied to was making. As an extreme example, imagine a billion data points that fit perfectly on a straight line. Then superimpose a million data points randomly on top…
I agree that a scatter plot is not the best for showing data with that many data points, but frankly it's kind of irrelevant. The point the author was making was just that it isn't that hard to cook up some data that is not highly correlated, but will be if you bin and average it.
Re: The 50 million dollar lie
#143" In New York City, the value-added score is actually not based on comparing the scores of a group of students from one year to the next, but on comparing the ‘predicted’ scores of a group of students to what those students actually get. The formula to generate this prediction is quite complicated, but the main piece of data it uses is the actual scores that the group of students got in the previous year. " http://ga…
You are correct that if students were assigned in some biased manner it could invalidate any relationship between teacher quality and subsequent student performance. That is exactly what makes this study interesting. All the students were assigned randomly! Much like in a randomized medical study this allows the establishment of a causal relationship.
Two thoughts...
1) Ethics: You put some children with dyslexia or with challenging behaviour with newly trained teachers with little experience? Did their parents have any say in this? Was there an ethics process? Was there support available for newly trained teachers meeting major needs in their randomly allocated classes? Amazing.
2) Correlation does not imply a causal relationship. You need a theory. Theory in this area is not positivist, it is fuzzier.
As I said before, 'good luck with that'.
Re: The 50 million dollar lie
#144Earlier quoted context omitted.
It seems to me that the evidence of weak correlation isn't quite as weak as you imply :-) If there's a dense line in there, how come it's not sticking out of the blob even a little? Do actual phenomena often yield "deceptive" blobby scatterplots with dense lines hidden inside them? That seems apriori improbable to me, thus increasing the Bayesian probability that the blobby scatterplot is not a misleading picture aft…
>If there's a dense line in there, how come it's not sticking out of the blob even a little? It is. If you look at the density in the upper right of the cloud, it is clearly clustered around the x=y line. The correlation may not be very strong, but without extra information about how strong of an effect would constitute clear evidence in favor of including the value add metric, we can't say anything more about it. >…
Um, that's not much to go on. Can you give some examples?
Re: The 50 million dollar lie
#145Earlier quoted context omitted.
For teachers, the metric is (roughly speaking) "% of students capable of multiplying/dividing numbers up to 4 digits in a standardized test setting". How do you game this metric? The fact that one company used a couple of bad objective metrics doesn't mean all objective metrics are bad. They are used with a great deal of success in many fields. Sales people are paid on commission, traders are paid proportionally to (…
One popular way of gaming the metric is systematic cheating on the tests by the teachers. I say popular because it has happened on a large scale, most recently in Atlanta as I recall. Another way to manipulate the test results is to manipulate which students are in your class or your school. The point is, people are endlessly creative in subverting rules to their own benefit - so they conform to the letter of the rul…
The fact that the current system has a bunch of cheaters is not an argument against more carefully and objectively measuring the current system. What next - bankers sometimes engage in rogue trading, so we should reduce monitoring of their behavior?
Another way to manipulate the test results is to manipulate which students are in your class or your school.
This is very difficult with VAM, since the goal is to increase (actual score - statistically predicted score). You need to reliably identify students who will do better than their statistical predictor.
I.e., you need to discover students who will improve drastically this year and then pack your student body with them.
Re: The 50 million dollar lie
#146Earlier quoted context omitted.
one can tell there is no line which most of the points lie near No, that's not correct. You can't conclude anything from looking at a big blob because you don't know the density of points at different places in the blob. This is the point the guy you replied to was making. As an extreme example, imagine a billion data points that fit perfectly on a straight line. Then superimpose a million data points randomly on top…
Okay hows this: unless the author of the original blog post is deliberately deceiving the audience but putting a bunch of points on top of each other then one can tell there is no line which most of the points lie near. I agree that a scatter plot is not the best for showing data with that many data points, but frankly it's kind of irrelevant. The point the author was making was just that it isn't that hard to cook u…
See the discussion here: http://news.ycombinator.com/item?id=4027337
I stand by the claim made in my blog post. Don't use scatterplots, use a density plot instead.
Incidentally, according to a comment the author made, the correlation is actually 0.3. That's far better than his graph suggests.
Re: The 50 million dollar lie
#147Earlier quoted context omitted.
Conspicuously missing from the various weighting schemes they compare is one with 100% classroom observations. It's not that conspicuously absent as they don't use 100% weighting for anything. That said, they do attempt to maximize ability to predict state test scores, and in that optimization classroom observation plays the smallest role of the three metrics (2-9%). Given they appear to have to done the analysis acr…
I do find it odd that teachers would rather be judged based on one or two people observing their classroom and ignoring actual student output, rather than raw numbers. As a developer I'd much rather be evaluated on some metric I could optimize for (I think some metric measuring feature value/bugs/fix rate/etc...), rather than just my manager watching me code/debug a couple of times per month. Okay, I really don't und…
I was a quant trader for a while. A considerable chunk of my salary was determined by profit, hardly unreasonable.
Measuring developers in other areas is difficult because of heterogeneous goals - last year I built a search engine, this year I'm statistically tracking user behavior. Hard to compare one to the other.
Education does not suffer this problem - last year a teacher taught 30 kids to read. This year she taught 28 kids to read. The goal is always maximizing the fraction of kids who can read.
Re: The 50 million dollar lie
#148Earlier quoted context omitted.
One alternative is to not use these flawed/complicated/expensive metrics. So replace a metric with 0.1 correlation with one having 0.0 correlation? If you are advocating that we should use the school principal's opinion rather than VAM, why do you believe opinion is superior? Do you have evidence that principal's opinion has a higher correlation with student outcomes than VAM?
So replace a metric with 0.1 correlation with one having 0.0 correlation? Two problems with this: 1. The other metric almost certainly doesn't have an 0.0 correlation, but we don't know that for sure since the data wasn't released for whatever reason. 2. The metric with 0.1 correlation (or whatever the number is)... keep in mind the context. What is the correlation with? Test scores, something that can be and is ofte…
The author is arguing that because the measurement is noisy, we should ignore it. This is silly - it just means a single class-year's point estimate is noisy, and multiple class-years must be combined to form an accurate estimate.
If you had 1/3 of your salary determined by the LOC you wrote or the number of bugs you closed, you would game the system and maximize your salary...
Indeed - if 1/3 of my salary was determined by the number of kids in my class who can read better than when they are predicted to read at this age, I'd definitely try to make sure that reading skills improved.
Conversely, if a stupid metric such as student/principal opinion were used, I'd focus on jokes and friendliness over education.
(Well, actually I didn't back when I taught. But that sure didn't help my student evaluations...)
Re: The 50 million dollar lie
#149Earlier quoted context omitted.
Can all the goals of a 3rd grade education system be reduced to a purely mechanical list of stuff? Yes, I would hope that a multi-million dollar enterprise can clearly define their goals. What are the goals? To get children to read individual words? Or to get children to read a sentence, and obtain meaning from it? I don't know off the top of my head whether the latter should be learned by 3rd grade. Ultimately setti…
> However, regardless of what the goal is, you still haven't given an way to game the system apart from "teach kids to read [words/sentences]". Sure I have. You ignore everything that is not tested. This gets you children that pass the tests. But it ignores all the other work that schools should be doing, and it reduces education to the worst, least inspiring, mechanical drudge work. > I don't know off the top of my…
Ignore non-goals and focus on goals?
But it ignores all the other work that schools should be doing...
Such as?
But this should be easy to discover, right?
No. Choosing your goals is about subjective value choices. If nonsense words are intrinsically valuable, they should be included, otherwise they should be excluded.
You are conflating the setting of goals with the method used to achieve them. If phonics is superior (I agree with you that it is), it will achieve higher scores. If teaching children nonsense words helps them understand real words, then teachers wishing to maximize their score will teach them.
Re: The 50 million dollar lie
#150Earlier quoted context omitted.
So replace a metric with 0.1 correlation with one having 0.0 correlation? Two problems with this: 1. The other metric almost certainly doesn't have an 0.0 correlation, but we don't know that for sure since the data wasn't released for whatever reason. 2. The metric with 0.1 correlation (or whatever the number is)... keep in mind the context. What is the correlation with? Test scores, something that can be and is ofte…
The 0.3 correlation described in the article is the correlation between a teacher's VAM score in a single class last year and this year. The author is arguing that because the measurement is noisy, we should ignore it. This is silly - it just means a single class-year's point estimate is noisy, and multiple class-years must be combined to form an accurate estimate. If you had 1/3 of your salary determined by the LOC…
Not quite..
Indeed - if 1/3 of my salary was determined by the number of kids in my class who can read better than when they are predicted to read at this age, I'd definitely try to make sure that reading skills improved.
Keep in mind that "score on a reading test" and "reading skills" are not the same things. For instance, one would be significantly improved by spending valuable class time teaching students tricks about how to succeed on the standardized reading test. It's similar to the incentives you'd get by determining salary by LOC - even if LOC correlates with performance, basing pay on it encourages all kinds of nonsense that doesn't actually benefit anyone.
So in that context, we have a metric that very loosely correlates with this clearly flawed marker of success. And we're supposed to spend ungodly sums of money implementing this strategy? Come on.
Conversely, if a stupid metric such as student/principal opinion were used, I'd focus on jokes and friendliness over education.
Well, they do propose to use student and principal evaluations..