Exploiting the most prominent AI agent benchmarks
151–160 of 175 posts
Re: Exploiting the most prominent AI agent benchmarks
#152The irony that this was very clearly written by an LLM, double negation always the simplest and clearest tell.
Re: Exploiting the most prominent AI agent benchmarks
#153Highly recommend this approach, saves us tons of eval time.
Re: Exploiting the most prominent AI agent benchmarks
#154Re: Exploiting the most prominent AI agent benchmarks
#155Re: Exploiting the most prominent AI agent benchmarks
#156Re: Exploiting the most prominent AI agent benchmarks
#157Earlier quoted context omitted.
> hopefully changes the way benchmarking is done The purpose of a system is what it does. AI companies want adcopy, not legitimate benchmarks. Even this very paper will be twisted into a means to that end. "Oooo, AI is exploiting our benchmarks. Scary alignment problem!!!one! Our AI is so good we can't contain it, INVEST NOW!"
I work at OpenAI and I really don't find this to be the case. We're pretty diligent about applying search blocklists, closing hacking loopholes, and reading model outputs to catch unanticipated hacks. If we wanted to, we could choose to close our eyes and plug our ears and report higher scores for Terminal-bench, SWE-bench, etc. that technically comply with the reference implementation but aren't aligned with real va…
Of course, but that's the difference between sins of commission and sins of omission. The question is what "pretty diligent" actually translates to in practice. How many people will encourage delays in a model release or post-training improvement waiting "for more thorough evaluation"? How many popularized AI results can you vouch for on this?
The zeitgeist is to celebrate bias for action, avoiding analysis paralysis and shipping things (esp. with conference driven research culture, even before we get into thorny questions of market dynamics), so even if we have a few pockets of meticulous excellence, the incentive structure pushes towards making the whole field rot.
Re: Exploiting the most prominent AI agent benchmarks
#158"No reasoning. No capability. Just exploitation of how the score is computed." The irony that this was very clearly written by an LLM, double negation always the simplest and clearest tell.
Re: Exploiting the most prominent AI agent benchmarks
#159Re: Exploiting the most prominent AI agent benchmarks
#160This is a phenomenal paper on exploits and hopefully changes the way benchmarking is done. From the paper: We achieved near-perfect scores on all of them without solving a single task. The exploits range from the embarrassingly simple (sending {} to FieldWorkArena) to the technically involved (trojanizing binary wrappers in Terminal-Bench), but they all share a common thread: the evaluation was not designed to resist…
>hopefully changes the way benchmarking is done. Yeah the path forward is simple: check if the solutions actually contain solutions. If they contain exploits then that entire result is discarded.