Live data from Hacker News

Exploiting the most prominent AI agent benchmarks

rdi.berkeley.edu

161–170 of 175 posts

Re: Exploiting the most prominent AI agent benchmarks

#161

Earlier quoted context omitted.

It does not ignore the word. It subverts it, and that's the point. It's the system equivalent of "death of the author", which states that omes a work is written, the authors intent loses relevance and the work must be examined on its own. The aurhors opinion or relationship to the work carries no more weight than any other persons. That's not "true" in any demonstrable sense, but it can be a useful form of analysis.…

This is not how people outside of cybernetics use POSWID. From context it does not appear to be how SlinkyOnStairs was using it either. I think it's also trying to be too cute. The first two definitions of purpose on Wiktionary[A]: 1. The end for which something is done, is made or exists. 2. Function, role. People (uselessly) talking about the purpose of a system are often referring to #1, while POSWID is using it t…

> From context it does not appear to be how SlinkyOnStairs was using it either.

The exact definition of "purpose" doesn't matter much here.

The particular version of the heuristic used here is that the stated purpose and the actual purpose often differ. POSIWID being the observation that the actual purpose is reflected by the outcomes of the system, because if that isn't the case the system gets changed.

Thus, the observation about AI benchmarks. AI companies have had years now to stop using unreliable benchmarks as advertising material. There's been years of piece after piece about the problems with these benchmarks. And yet the AI marketing continues as is.

Re: Exploiting the most prominent AI agent benchmarks

#164

Earlier quoted context omitted.

I'm surprised anyone took them seriously in the first place.

We need good benchmarks or we are just left following the hype train.

The benchmarks are the hype train that’s what I’m saying.

Re: Exploiting the most prominent AI agent benchmarks

#168

Earlier quoted context omitted.

> hopefully changes the way benchmarking is done The purpose of a system is what it does. AI companies want adcopy, not legitimate benchmarks. Even this very paper will be twisted into a means to that end. "Oooo, AI is exploiting our benchmarks. Scary alignment problem!!!one! Our AI is so good we can't contain it, INVEST NOW!"

I work at OpenAI and I really don't find this to be the case. We're pretty diligent about applying search blocklists, closing hacking loopholes, and reading model outputs to catch unanticipated hacks. If we wanted to, we could choose to close our eyes and plug our ears and report higher scores for Terminal-bench, SWE-bench, etc. that technically comply with the reference implementation but aren't aligned with real va…

I work at runloop and I've spent a considerable amount of time getting various benchmarks to run with very high concurrency (thousands at once). My experience is similar to your own: it takes a ton of time and effort setting up benchmarks to run at scale with protection against reward hacks.

Keeping a benchmark test harness secure and fast is non-trivial. You need to keep the grading script and the solution off the box, use network controls, deal with external resource usage, etc. It's a lot of work. I don't think it's realistic to expect benchmark authors to bullet proof their benchmark runners. Most benchmarks are written to be run conveniently on a single machine (ie. in docker), not to run in parallel across tends of thousands of secure, isolated machines.

Re: Exploiting the most prominent AI agent benchmarks

#169

Earlier quoted context omitted.

This is not how people outside of cybernetics use POSWID. From context it does not appear to be how SlinkyOnStairs was using it either. I think it's also trying to be too cute. The first two definitions of purpose on Wiktionary[A]: 1. The end for which something is done, is made or exists. 2. Function, role. People (uselessly) talking about the purpose of a system are often referring to #1, while POSWID is using it t…

> From context it does not appear to be how SlinkyOnStairs was using it either. The exact definition of "purpose" doesn't matter much here. The particular version of the heuristic used here is that the stated purpose and the actual purpose often differ. POSIWID being the observation that the actual purpose is reflected by the outcomes of the system, because if that isn't the case the system gets changed . Thus, the o…

> POSIWID being the observation that the actual purpose is reflected by the outcomes of the system, because if that isn't the case the system gets changed.

I fundamentally disagree with this, and it seems to differ from how other proponents of POSIWID in this thread view POSIWID.

It also seems trivially false; systems are dynamic what was the purpose of the system just before it was changed because people didn't like the outcomes?

Re: Exploiting the most prominent AI agent benchmarks

#170
post #93

Earlier quoted context omitted.

I will propose that you are wrong. 1. We must ignore the intentions of the designers (your claim), and instead see what the outcomes are 2. Therefore we should ignore Beer's intentions when designing the phrase POSWID, and instead see how it is used. 3. The overwhelming majority of people using it on the internet (including the GP comment) is to imply that the people perpetuating the system actually desire the outcom…

There are plenty of cases where you absolutely can/should discuss outcomes in a way where the intention is not factored in because it can often be straight up irrelevant. If a gun is developed with the intention of hunting only bears and someone uses it to shoot people, you don’t have to constantly preface things by talking about how it’s supposed to be used only on bears. Sometimes that fact, depending on the contex…

I agree with the idea that intent is often irrelevant. I disagree that POSIWID is a good way to communicate that idea.
Post reply on HN