Goodhart's Law Comes for Every Benchmark You Trust
21–30 of 50 posts
Re: Goodhart's Law Comes for Every Benchmark You Trust
#22Had to stop reading when the article devolved into Claude spam. "defensible in isolation," "honestly ranked," ugh. Please write your own blog post.
My strategy these days is to scan and look for the tells and click out when I see them. Mine was the same "honestly ranked, with no silver bullets on offer". I suspect in less than a year we won't be able to tell the difference.
So why haven’t they? My theory is they see this as a sort of fingerprint, useful to not train on later. Or something. Maybe they just don’t care. Certainly feels either intentional or a result of ambivalence.
It’s certainly true today that I probably wouldn’t know an AI written article if the author went out of their way to use one of the many prompts available to tone down the AI-isms.
Re: Goodhart's Law Comes for Every Benchmark You Trust
#23I view these AI benchmarks the same. No I do not care that GPT got 1200 on FartAGIMaX-4.0-Extreme and Claude got 1350. I care about how much it costs and how correctly it does the tasks that I give it. Unfortunately the only way to know is to use them all myself and measure it myself.
At the end of the day these things are all so damn close in how they behave in whatever harness so it realy just does boil down to whatever is actually cheapest.
This is why Deepseek is great: it’s so much cheaper it doesn’t matter if I burn way more tokens because it’s still orders of magnitude cheaper than the US SotA models. If it doesn’t get it quite right immediately I just do a few more turns and then it’s fine. Barely an inconvenience.
Re: Goodhart's Law Comes for Every Benchmark You Trust
#24Had to stop reading when the article devolved into Claude spam. "defensible in isolation," "honestly ranked," ugh. Please write your own blog post.
I often read as much as 1000 words thinking to myself: “this is smooth”, but then think to myself “Is this my friend Opus 4.8 or now Opus 5?.
In this case the rhetorical neatness is unmistakable especially in the beginnings and endings of paragraphs: “Here is the constructive turn, honestly ranked, with no silver bullets on offer.”
Yes: and that is actually a smoking gun.
Re: Goodhart's Law Comes for Every Benchmark You Trust
#25Had to stop reading when the article devolved into Claude spam. "defensible in isolation," "honestly ranked," ugh. Please write your own blog post.
My strategy these days is to scan and look for the tells and click out when I see them. Mine was the same "honestly ranked, with no silver bullets on offer". I suspect in less than a year we won't be able to tell the difference.
I think you are right that in a year or two LLMs will be able to do a good impression of many technical styles. But not Nabokov, Kundera, or Kafka for subtlety.
If one manages to channel Edsger W. Dijkstra I will be impressed and rank it high on my leaderboard.
Re: Goodhart's Law Comes for Every Benchmark You Trust
#26Earlier quoted context omitted.
My strategy these days is to scan and look for the tells and click out when I see them. Mine was the same "honestly ranked, with no silver bullets on offer". I suspect in less than a year we won't be able to tell the difference.
I’ve thought this for a while, but why hasn’t it happened yet? At this point, OpenAI and Anthropic and friends could definitely remove the AI “smell” from writing output, or give users a first class way to specify a writing style. So why haven’t they? My theory is they see this as a sort of fingerprint, useful to not train on later. Or something. Maybe they just don’t care. Certainly feels either intentional or a res…
Re: Goodhart's Law Comes for Every Benchmark You Trust
#27Earlier quoted context omitted.
My strategy these days is to scan and look for the tells and click out when I see them. Mine was the same "honestly ranked, with no silver bullets on offer". I suspect in less than a year we won't be able to tell the difference.
I’ve thought this for a while, but why hasn’t it happened yet? At this point, OpenAI and Anthropic and friends could definitely remove the AI “smell” from writing output, or give users a first class way to specify a writing style. So why haven’t they? My theory is they see this as a sort of fingerprint, useful to not train on later. Or something. Maybe they just don’t care. Certainly feels either intentional or a res…
Re: Goodhart's Law Comes for Every Benchmark You Trust
#28Earlier quoted context omitted.
My strategy these days is to scan and look for the tells and click out when I see them. Mine was the same "honestly ranked, with no silver bullets on offer". I suspect in less than a year we won't be able to tell the difference.
In a year or more, the audience will likely be an agent instead of a person. It'll be interesting to see how that shifts language and article formats.
Re: Goodhart's Law Comes for Every Benchmark You Trust
#29Jokes on them, I don't trust benchmarks Once something becomes a benchmark it is no longer a good benchmark.
No, it's when a benchmark becomes a target. You might have a private benchmark that you tell no one about. Would you not trust it?
Re: Goodhart's Law Comes for Every Benchmark You Trust
#30Earlier quoted context omitted.
In a year or more, the audience will likely be an agent instead of a person. It'll be interesting to see how that shifts language and article formats.
It has been for years. The actual audience for a great deal of text you see is Google’s ranking system. Look up any recipe and ask yourself who reads the ten paragraph story time about grandma’s cookies. It literally isn’t intended to be read.