Live data from Hacker News

Goodhart's Law Comes for Every Benchmark You Trust

cacm.acm.org

31–40 of 50 posts

Re: Goodhart's Law Comes for Every Benchmark You Trust

#31

Earlier quoted context omitted.

No, it's when a benchmark becomes a target. You might have a private benchmark that you tell no one about. Would you not trust it?

Good benchmarks are costly to build even for mid-large corporations. And once the benchmark is used on models you really can’t tell if the problems would be scrapped for training

I was speaking in general, not just about AI.

Re: Goodhart's Law Comes for Every Benchmark You Trust

#32
post #11

Clearly, the solution is to judge society by how many currently-un-gamed benchmarks it has produced.

Benchmark is by definition gamed. That is the essence of Goodhart's law.

Society is a benchmark of sorts. You see where that leads.

---- edit: as in this quote from Gulliver:

I told him, “that in the kingdom of Tribnia, by the natives called Langdon, where I had sojourned some time in my travels, the bulk of the people consist in a manner wholly of discoverers, witnesses, informers, accusers, prosecutors, evidences, swearers, together with their several subservient and subaltern instruments, all under the colours, the conduct, and the pay of ministers of state, and their deputies. The plots, in that kingdom, are usually the workmanship of those persons who desire to raise their own characters of profound politicians; to restore new vigour to a crazy administration; to stifle or divert general discontents; to fill their coffers with forfeitures; and raise, or sink the opinion of public credit, as either shall best answer their private advantage. It is first agreed and settled among them, what suspected persons shall be accused of a plot; then, effectual care is taken to secure all their letters and papers, and put the owners in chains. These papers are delivered to a set of artists, very dexterous in finding out the mysterious meanings of words, syllables, and letters: for instance, they can discover a close stool, to signify a privy council; a flock of geese, a senate; a lame dog, an invader; the plague, a standing army; a buzzard, a prime minister; the gout, a high priest; a gibbet, a secretary of state; a chamber pot, a committee of grandees; a sieve, a court lady; a broom, a revolution; a mouse-trap, an employment; a bottomless pit, a treasury; a sink, a court; a cap and bells, a favourite; a broken reed, a court of justice; an empty tun, a general; a running sore, the administration.

When this method fails, they have two others more effectual, which the learned among them call acrostics and anagrams. First, they can decipher all initial letters into political meanings. Thus N, shall signify a plot; B, a regiment of horse; L, a fleet at sea; or, secondly, by transposing the letters of the alphabet in any suspected paper, they can lay open the deepest designs of a discontented party. So, for example, if I should say, in a letter to a friend, ‘Our brother Tom has just got the piles,’ a skilful decipherer would discover, that the same letters which compose that sentence, may be analysed into the following words, ‘Resist -, a plot is brought home - The tour.’ And this is the anagrammatic method.”

== some ecclesiastes quote would be nice here.

Re: Goodhart's Law Comes for Every Benchmark You Trust

#33

Had to stop reading when the article devolved into Claude spam. "defensible in isolation," "honestly ranked," ugh. Please write your own blog post.

This is the second ACM article in the last few months that was clearly not written by a human. Funny because it's against their editorial policy... what is going on over there?

Re: Goodhart's Law Comes for Every Benchmark You Trust

#34
post #16

Earlier quoted context omitted.

My strategy these days is to scan and look for the tells and click out when I see them. Mine was the same "honestly ranked, with no silver bullets on offer". I suspect in less than a year we won't be able to tell the difference.

I’ve thought this for a while, but why hasn’t it happened yet? At this point, OpenAI and Anthropic and friends could definitely remove the AI “smell” from writing output, or give users a first class way to specify a writing style. So why haven’t they? My theory is they see this as a sort of fingerprint, useful to not train on later. Or something. Maybe they just don’t care. Certainly feels either intentional or a res…

If the underlying prompt of the model stays the same, it seems to me that LLMs will always have common tells unless overridden with a thorough prompt from the end user. It's like if you had 1 person write half of the content on the internet. You'd probably get pretty good at noticing their writing style.

Maybe my understanding of LLMs is wrong, but it seems obvious to me that when you have a large corpus of LLM output you will eventually notice common tells when everyone is using the same models, weights, and base prompt.

Re: Goodhart's Law Comes for Every Benchmark You Trust

#36
post #16

Had to stop reading when the article devolved into Claude spam. "defensible in isolation," "honestly ranked," ugh. Please write your own blog post.

My strategy these days is to scan and look for the tells and click out when I see them. Mine was the same "honestly ranked, with no silver bullets on offer". I suspect in less than a year we won't be able to tell the difference.

This article was too verbose for me to want to read it, but I'm still not sure about it being LLM-gen'd. All I got is Claude says "honest" a lot.

Re: Goodhart's Law Comes for Every Benchmark You Trust

#38
The obvious solution is to have non-public benchmarks.

It is exceedingly difficult to train on a proprietary benchmark administered by someone with half a brain (i.e. don't sign up for a ChatGPT account with your benchmark@artificialanalysis.ai email) - you have to find a tiny needle in a vast haystack.

In fact, it can be difficult enough that it's simply not economically viable - that is, that it's cheaper to make the model better than it is to try to find the account running the benchmark.

In the limit case, the benchmark is indistinguishable from...normal problems that need to be solved.

Re: Goodhart's Law Comes for Every Benchmark You Trust

#40
post #32
post #11

Clearly, the solution is to judge society by how many currently-un-gamed benchmarks it has produced.

Benchmark is by definition gamed. That is the essence of Goodhart's law. Society is a benchmark of sorts. You see where that leads. ---- edit: as in this quote from Gulliver: I told him, “that in the kingdom of Tribnia, by the natives called Langdon, where I had sojourned some time in my travels, the bulk of the people consist in a manner wholly of discoverers, witnesses, informers, accusers, prosecutors, evidences,…

I don't understand your point, or your quote.

Yes, if you try hard enough, you will find proof for any accusation whose outcome you've already decided. But what does that have to do with anything about society being a benchmark?

Post reply on HN