I have spoken to team members, and they all say the results of this and coder are very, very much leakage (no suprisse given the result!!)
yea, that's my first thought seeing the result too. we need a reputable proprietary eval.
I think this is self-conflicting. If the evaluation is proprietary then it is most certainly not reputable. We'd want open metrics where we can analyze the limitations. Of course, we'd need open data too, but that's exceptionally rare these days. Plus, a metric isn't going to really tell us if we have have spoilage or not. You can get some evidence for spoilage through a trained model, but it is less direct, fuzzier, and more tells us about what information it was able to memorize rather than if the data was spoiled.