So if the teacher evaluation method is alleged to be inferior, please provide a new evaluation method which is superior to it. Can't? Won't? Then be quiet.
The idea would be not to waste numerous resources and money implementing a broken idea and spend some time coming up with a better one. Why put the education of our children and the careers of good teachers on the line with such a poorly thought out and half-assed idea?
The 50 million dollar lie
51–60 of 159 posts
Re: The 50 million dollar lie
#52Re: The 50 million dollar lie
#53http://news.ycombinator.com/item?id=4728123
and I'm glad to see that so many participants, from the founder on to the newest member, enjoy thinking about and checking facts on education issues.
There is other research on the issue of teacher quality and value-added measures to assess teacher quality,
http://hanushek.stanford.edu/publications/teacher-quality
and I encourage Hacker News participants to dig through some of the other publications to figure out what the best available current data show.
The other comments here suggest that the blog author of the submitted blog post is overstating his own case at least as much as he accuses the Gates Foundation supporters of overstating their case. If people who oppose the kind of teacher ratings proposed by the Gates Foundation really had courage of their convictions, they might try to let learners have full power to shop for teachers, on the grounds that teacher quality measures miss many dimensions of teacher performance that are important to individual learners. In fact, very few states even have as easy and time-proven a policy reform as statewide open enrollment (which my state has had for almost a generation now)
http://education.state.mn.us/MDE/StuSuc/EnrollChoice/index.h...
and no state yet lets learners shop for schools on an equal per-capita funding basis whether the school is state-operated or not (as the Netherlands has for more than a century). Teachers will have plenty of deserved professional respect and regard from their clients if their clients have power to shop. But if policy proposals for learner power to shop are rejected, hand in hand with rejecting proposals for evaluating teachers for their effectiveness in publicly subsidized instruction, I wonder what the true agenda is. What assurance do we (we parents, we taxpayers, we members of the general public) have that all schools are staffed with the best available teachers unless someone is checking on how the teachers are doing their work?
Re: The 50 million dollar lie
#54I know nothing about the domain of teacher measurement, and have no opinion on it. But the first chart in the blog post, with the big mass of blue on it, is surely not the strong evidence of weak correlation, that the author is making it out to be? There could still be a strong correlation in that data, even if it looks like a blob of blue - because there are so many data points on that chart, that we can no longer t…
I will make a strong claim: If doctor effectiveness were this stable, and it were the only data I had available, I would choose a doctor on this metric, ceteris paribus.
I found a scatterplot by the author that uses smaller marks so it looks a lot less blob-ish and more heat-map-ish:
http://garyrubinstein.teachforus.org/2012/02/26/analyzing-re...
I'm trying to track down the original data to verify this myself, but a couple of quick points:
There definitely seems to be good evidence of a relationship. It might be "weak" but that's not the same thing as "ineffective". Considering the number of data points, it makes it unlikely that the trend occurred by chance. This plot also makes it clear why the averaging made it look so pretty: the average in just about any percentile group in 08-09 matched 09-10 regardless of how finely it's divided (to a point).
It's certainly the case that there's a lot of "noise". The question is how to interpret it. One might conclude that there is a lot of "measurement noise" -- that teacher evaluations are inherently messy and rather uninformative except as a larger trend. There seems to be reasonable support for this! Students aren't the same, even from year to year. Of course it could "signal noise" - maybe teachers themselves change from year to year. Perhaps one year you're more motivated, the next your not, or vice versa.
To understand the implications of this noise, we also need to ask what we're using it for. For instance, if this were a plot of the correlation between money donated to Watsi and quality-of-life, you'd probably think you were doing pretty well! You might point out that a particular quality-of-life metric is flawed or noisy, or that the quality-of-life depends on lots of things besides health, but likely you'd be happy that you're clearly causing some benefit.
So when we ask whether the data suggest the viability of value-added measurements in aggregate, to me that seems to be a yes. But if we ask whether the data suggest using these measurements to make a decision on a per-teacher basis, that depends on the context of the alternatives! If this is the only thing we have to go by, perhaps it's not so bad.
I would love to see how stable the other parts of the teacher evaluation fit in. If you combine multiple noisy signals, you often get a much better picture of what's going on! This accounts for only 1/3 of the teacher's evaluation. I would like to think that if you appeared not to be a value-added teacher, but your principals and coworkers spoke highly of you, that you'd still get that raise.
Also, I'm looking forward to multiple years of this data - if your score accounts for 35 percent of the variance from one year to the next, multiple years might look a lot better. Perhaps in deciding teacher tenure, scores over 5-10 years could be really useful.
Re: The 50 million dollar lie
#55To build intuition, let's consider a hypothetical society. We divide a group of people in two. The first group rolls a die with 100 sides labeled from 1 to 100 in even increments. The second group rolls a die with labeled 1 to 110 in even increments. The roll determines their annual ability level which is unknowable. If you were hiring, which group would you prefer? Is it fair to the second group? Note that this isn'…
To provide a contrary hypothetical... Imagine if one group had 55% of its dice labeled up to 110 and 45% up to 100, and the other group had the opposite. So if all you use for hiring is the group membership, you're going to very arbitrarily miss out on a lot of 110s and hire a lot of 100s. And now also imagine that this process of identifying who is in what group is ridiculously expensive.
Re: The 50 million dollar lie
#56I suspect there'll be a lot of differing opinions on this, and I'm looking forward to seeing the discussion. What would be helpful would be if people could say how much experience they have in hard statistics, and how much what they say is driven a priori from the data. I know that hackers, in particular, have real problems with "Argument from Authority", but stats is one place where it's really, really easy to go wr…
I'll bite. I won't discuss my experience in my posts since I don't want to argue from authority but will do so below. But first, I should point out that you're committing a similar sin to that which is alleged in the article. What you really care about is the correctness of an argument. You hypothesize that formal training is an important indicator of correctness. Presumably there's also noise in that scatterplot but…
If you have no such evidence then the onus will be on you to make your argument more complete, more coherent, and more comprehensive. If you have evidence (note: evidence, not proof) then you can be a little less rigorous in what you say, and rely on people giving you the benefit of the doubt while they work through the argument.
What I see a lot of is long, apparently good arguments, that then turn out not to be as complete or coherent. they sometimes just don't hang together.
Significant amounts of formal study in a subject is evidence that someone might just have a better understanding. After spending a lot of time on the internet I'm tired of having to wade through every single argument in detail looking for all the possible chinks.
Maybe that's just impossible. Maybe every person has to redo every analysis for every argument. Seems like a complete waste of almost everyone's time. What about "Don't Repeat Yourself" or "Don't re-invent the wheel." I guess we are doomed to reinvent the wheel in every single discussion.
Re: The 50 million dollar lie
#57This is a surprising blog post in that I draw the complete opposite conclusion the author does. The author seems to think that the averaging hides volatility (which it does), which leads to incorrect conclusions drawn. Whereas to me it looks like it removes visual noise to show an actual trend. In his original scatter plot, because of the big ball in the middle it's easy to handwave and say, "look a random blob" -- b…
"Whereas to me it looks like it removes visual noise to show an actual trend." Except for that in this case the noise is more important than the trend. Think about it, if you're firing or sanctioning perhaps 30%+ of teachers each year for no reason, then only complete morons would go into teaching. It's the same as airport security, where a .1% false positive rate is unacceptable, whereas a 75% false negative rate is…
Re: The 50 million dollar lie
#58Earlier quoted context omitted.
The idea would be not to waste numerous resources and money implementing a broken idea and spend some time coming up with a better one. Why put the education of our children and the careers of good teachers on the line with such a poorly thought out and half-assed idea?
If you think that no time has been spent coming up with this idea, you are coming very late to the party. This debate is decades old, at a minimum. Dates from 1971, says Wikipedia. http://en.wikipedia.org/wiki/Value-added_modeling It's not something that "Bill Gates" or anyone else made up yesterday. It's a state-of-the-art approach to a difficult problem. If you have a better way, let's hear it. Otherwise, be quiet.
Re: The 50 million dollar lie
#59There is no objective way to measure teacher performance. Any evaluation method that can be written as a list of rules can and will quickly be gamed. The thing is, it's easy to find out who the best teachers are. Simply ask the students, parents, and staff. They all know who the good ones and the bad ones are - and that can't be gamed.
Are you claiming that there's something special about teaching that makes teacher performance much more difficult, or even impossible, to quantify? If so, what is the point of teaching if one cannot measure results?
Your solution to "ask the students, parents, and staff" is precisely a measurement method (albeit a more qualitative than quantitative one), and moreover, one that can be gamed easily by anyone who's socially shrewd. Qualitative measures like that is exactly why we have so many crappy politicians.
Re: The 50 million dollar lie
#60So if the teacher evaluation method is alleged to be inferior, please provide a new evaluation method which is superior to it. Can't? Won't? Then be quiet.
In science, falsifying a theory does not require that you replace it with a better theory.