Both sides are right but they seem to be concerned with different things.
Grouping data points does mask volatility. This doesn't make the effect any less real on the societal level. If we seek to improve student performance, skewing the pool of teachers towards the better ones will do that.
Yet there is also an enormous amount of volatility which is also real. Ideally we would just create better predictors. Constraining ourself for the moment to this set, since they have high error rates on the individual level, it's also reasonable to assert that in relying on this metric, we will make unfair decisions for many teachers.
The right solution depends on how much you value optimizing student outcomes vs. optimizing fairness as well as second order effects such as driving away teachers. My bias would be to use these scores for economic incentives to attempt to ensure that retention for the best teachers is higher than for the worst. Without additional work, termination likely isn't warranted based on this data.
To those who argue that pay differentials require better evidence, I would suggest comparing the validity of performance evaluations across any other job. It's imperfect but better for society than doing nothing.
For long-term fairness, it's important that there isn't a single value-added model. Competitive pressure can do wonders to improve model quality and it's also fairer since those who don't do well on one model can move to a different rubric.