Live data from Hacker News

Exploiting the most prominent AI agent benchmarks

rdi.berkeley.edu

101–110 of 175 posts

Re: Exploiting the most prominent AI agent benchmarks

#102

Earlier quoted context omitted.

What was the cheat in the 2024 Intel situation? The TomsHardware article and the Phoronix article they linked were quite vague. (Not to say I have any doubts, just curious, hadn’t heard of this one).

Intel basically benchmaxxed their compiler optimizations. They used detailed knowledge of the benchmark to make their compiler generate machine code to do better on the benchmark in a way that was not beneficial for non-benchmark scenarios.

I assumed as much, I’m just wondering what exactly they did. For example IIRC some phone company would detect that a benchmark was running by checking for the program name, and then allow the clock to boost higher (increase thermal limits) if it was a benchmark (like you could literally avoid the cheating behavior by changing the name of the program being run).

Re: Exploiting the most prominent AI agent benchmarks

#103

Earlier quoted context omitted.

https://en.wikipedia.org/wiki/The_purpose_of_a_system_is_wha... You are misunderstanding the saying. It is entirely about unintended consequences and viewing the system for what it actually does and not any stated intentions of the designers.

Well that’s stupid and completely ignores the meaning of the word “purpose”.

It does not ignore the word. It subverts it, and that's the point. It's the system equivalent of "death of the author", which states that omes a work is written, the authors intent loses relevance and the work must be examined on its own. The aurhors opinion or relationship to the work carries no more weight than any other persons.

That's not "true" in any demonstrable sense, but it can be a useful form of analysis. As it is with "purpose of a system"

Re: Exploiting the most prominent AI agent benchmarks

#104
post #63
post #17

If only the blog itself wasn't written by AI? >No reasoning. No capability. Just exploitation of how the score is computed. shudder

Agreed. The premise is interesting but reading content like this is grating.

im actually getting so tilted that people can't just be forthcoming about when they used AI to write something. 99% of readme.mds i run into now on github piss me off. out of all the things people could cede to automation, they foolishly went and self-owned their ability to communicate. smfh.

if you've worked on something diligently and understand it and have novel insight to share, let's hear _your_ damn voice.

Re: Exploiting the most prominent AI agent benchmarks

#106

Earlier quoted context omitted.

> hopefully changes the way benchmarking is done The purpose of a system is what it does. AI companies want adcopy, not legitimate benchmarks. Even this very paper will be twisted into a means to that end. "Oooo, AI is exploiting our benchmarks. Scary alignment problem!!!one! Our AI is so good we can't contain it, INVEST NOW!"

>The purpose of a system is what it does. I am so tired of this saying. It's not true, in general. Systems almost universally have unintended consequences and result in side effects their designers did not foresee. Designing benchmarks resistant to adversarial attempts to exploit the benchmark software is just something no one was thinking about when they created SWE-bench.

I think the point of the saying is that as systems tend to expand, sooner or later we become part of them. That means that we can no longer see them from outside, we're now part of the system and our goals and the system's goals will align. Then the purpose of the system can't be anything else than what it does.

Re: Exploiting the most prominent AI agent benchmarks

#109
post #2

This is a phenomenal paper on exploits and hopefully changes the way benchmarking is done. From the paper: We achieved near-perfect scores on all of them without solving a single task. The exploits range from the embarrassingly simple (sending {} to FieldWorkArena) to the technically involved (trojanizing binary wrappers in Terminal-Bench), but they all share a common thread: the evaluation was not designed to resist…

>hopefully changes the way benchmarking is done. Yeah the path forward is simple: check if the solutions actually contain solutions. If they contain exploits then that entire result is discarded.

But that requires me to do things :(
Post reply on HN