what are the point of benchmarks?
Are you serious? To help you pick a model.
Exploiting the most prominent AI agent benchmarks
101–110 of 175 posts
Re: Exploiting the most prominent AI agent benchmarks
#102Earlier quoted context omitted.
What was the cheat in the 2024 Intel situation? The TomsHardware article and the Phoronix article they linked were quite vague. (Not to say I have any doubts, just curious, hadn’t heard of this one).
Intel basically benchmaxxed their compiler optimizations. They used detailed knowledge of the benchmark to make their compiler generate machine code to do better on the benchmark in a way that was not beneficial for non-benchmark scenarios.
Re: Exploiting the most prominent AI agent benchmarks
#103Earlier quoted context omitted.
https://en.wikipedia.org/wiki/The_purpose_of_a_system_is_wha... You are misunderstanding the saying. It is entirely about unintended consequences and viewing the system for what it actually does and not any stated intentions of the designers.
Well that’s stupid and completely ignores the meaning of the word “purpose”.
That's not "true" in any demonstrable sense, but it can be a useful form of analysis. As it is with "purpose of a system"
Re: Exploiting the most prominent AI agent benchmarks
#104If only the blog itself wasn't written by AI? >No reasoning. No capability. Just exploitation of how the score is computed. shudder
Agreed. The premise is interesting but reading content like this is grating.
if you've worked on something diligently and understand it and have novel insight to share, let's hear _your_ damn voice.
Re: Exploiting the most prominent AI agent benchmarks
#105Re: Exploiting the most prominent AI agent benchmarks
#106Earlier quoted context omitted.
> hopefully changes the way benchmarking is done The purpose of a system is what it does. AI companies want adcopy, not legitimate benchmarks. Even this very paper will be twisted into a means to that end. "Oooo, AI is exploiting our benchmarks. Scary alignment problem!!!one! Our AI is so good we can't contain it, INVEST NOW!"
>The purpose of a system is what it does. I am so tired of this saying. It's not true, in general. Systems almost universally have unintended consequences and result in side effects their designers did not foresee. Designing benchmarks resistant to adversarial attempts to exploit the benchmark software is just something no one was thinking about when they created SWE-bench.
Re: Exploiting the most prominent AI agent benchmarks
#107Re: Exploiting the most prominent AI agent benchmarks
#108Re: Exploiting the most prominent AI agent benchmarks
#109This is a phenomenal paper on exploits and hopefully changes the way benchmarking is done. From the paper: We achieved near-perfect scores on all of them without solving a single task. The exploits range from the embarrassingly simple (sending {} to FieldWorkArena) to the technically involved (trojanizing binary wrappers in Terminal-Bench), but they all share a common thread: the evaluation was not designed to resist…
>hopefully changes the way benchmarking is done. Yeah the path forward is simple: check if the solutions actually contain solutions. If they contain exploits then that entire result is discarded.