Clearly, the solution is to judge society by how many currently-un-gamed benchmarks it has produced.
Goodhart's Law Comes for Every Benchmark You Trust
41–50 of 50 posts
Re: Goodhart's Law Comes for Every Benchmark You Trust
#42Earlier quoted context omitted.
My strategy these days is to scan and look for the tells and click out when I see them. Mine was the same "honestly ranked, with no silver bullets on offer". I suspect in less than a year we won't be able to tell the difference.
I’ve thought this for a while, but why hasn’t it happened yet? At this point, OpenAI and Anthropic and friends could definitely remove the AI “smell” from writing output, or give users a first class way to specify a writing style. So why haven’t they? My theory is they see this as a sort of fingerprint, useful to not train on later. Or something. Maybe they just don’t care. Certainly feels either intentional or a res…
This wiki page is updated frequently and only needs to be fed into an AI to remove its telltale writing.
Re: Goodhart's Law Comes for Every Benchmark You Trust
#43Had to stop reading when the article devolved into Claude spam. "defensible in isolation," "honestly ranked," ugh. Please write your own blog post.
Re: Goodhart's Law Comes for Every Benchmark You Trust
#44A model can crush some test and still be a pain to use in real life.
Re: Goodhart's Law Comes for Every Benchmark You Trust
#45Re: Goodhart's Law Comes for Every Benchmark You Trust
#46Re: Goodhart's Law Comes for Every Benchmark You Trust
#47Broken link?
Re: Goodhart's Law Comes for Every Benchmark You Trust
#48Earlier quoted context omitted.
Benchmark is by definition gamed. That is the essence of Goodhart's law. Society is a benchmark of sorts. You see where that leads. ---- edit: as in this quote from Gulliver: I told him, “that in the kingdom of Tribnia, by the natives called Langdon, where I had sojourned some time in my travels, the bulk of the people consist in a manner wholly of discoverers, witnesses, informers, accusers, prosecutors, evidences,…
I don't understand your point, or your quote. Yes, if you try hard enough, you will find proof for any accusation whose outcome you've already decided. But what does that have to do with anything about society being a benchmark?
Re: Goodhart's Law Comes for Every Benchmark You Trust
#49obviously, the best benchmark is the one you tell no one about.
Kind of like when you think you come up with what you think is the perfect question that can't be misinterpreted and then you ask someone with a case of weaponized autism about it and they show a myriad of different ways a person can parse what you're trying to do.
This is why some certifications don't ask you for 'correct' answers in the sense of a logically complete answer that fulfills the condition, they ask for something like "The Cisco Way". Where the correct answer is how they teach it in the book so anyone that works on routers or whatever does it the same way.
Re: Goodhart's Law Comes for Every Benchmark You Trust
#50Earlier quoted context omitted.
Good benchmarks are costly to build even for mid-large corporations. And once the benchmark is used on models you really can’t tell if the problems would be scrapped for training
I was speaking in general, not just about AI.