If only the blog itself wasn't written by AI? >No reasoning. No capability. Just exploitation of how the score is computed. shudder
Exploiting the most prominent AI agent benchmarks
31–40 of 175 posts
Re: Exploiting the most prominent AI agent benchmarks
#32If only the blog itself wasn't written by AI? >No reasoning. No capability. Just exploitation of how the score is computed. shudder
Yes, marks of AI all over the place. Also the SVGs. >No solution written, 100% score. Its weird. Turns out that hardest problem for LLMs to really tackle is long-form text.
Re: Exploiting the most prominent AI agent benchmarks
#33I think we should all consider the possibility that part of the reason Anthropic hasn't immediately released Mythos is that it would be slightly disappointing relative to the benchmark scores.
I’m convinced specialised models are the way but this means writing off the investment in existing assets which they won’t do for obvious reasons.
Re: Exploiting the most prominent AI agent benchmarks
#34This team is doing a good job. They use problems that were created in last 30days to avoid training set leakage. https://swe-rebench.com/
Re: Exploiting the most prominent AI agent benchmarks
#35what are the point of benchmarks?
Re: Exploiting the most prominent AI agent benchmarks
#36Re: Exploiting the most prominent AI agent benchmarks
#37Earlier quoted context omitted.
>hopefully changes the way benchmarking is done. Yeah the path forward is simple: check if the solutions actually contain solutions. If they contain exploits then that entire result is discarded.
In human multiple choice tests they sometimes use negative marking to discourage guessing. It feels like exploits should cancel out several correct solutions.
The Artificial Analysis Omniscience benchmark does penalize guessing, so it actually helps you determine which LLMs are likely to just guess rather than telling you they don't know. Only a very few of the frontier models actually score higher than 0 on this, where 0 means that it's equally likely to return a correct answer as it is to return a hallucination on factual questions.
Re: Exploiting the most prominent AI agent benchmarks
#38Evaluating AI models has always relied largely on trust. If you want to game the benchmarks, you can. Simply train on your test data.
When an AI agent has autonomous control over the same computing environment where its scores are recorded, it's not surprising that it can, in principle, falsify its scores. A more interesting question would be whether agents behave in this way automatically, without manual tuning by the researcher.
That said, the main takeaway of "don't trust the number, trust the methodology" is valid. It's already a truism for researchers, and spreading the word to non-researchers is valuable.
Re: Exploiting the most prominent AI agent benchmarks
#39Re: Exploiting the most prominent AI agent benchmarks
#40This is a phenomenal paper on exploits and hopefully changes the way benchmarking is done. From the paper: We achieved near-perfect scores on all of them without solving a single task. The exploits range from the embarrassingly simple (sending {} to FieldWorkArena) to the technically involved (trojanizing binary wrappers in Terminal-Bench), but they all share a common thread: the evaluation was not designed to resist…
The purpose of a system is what it does.
AI companies want adcopy, not legitimate benchmarks. Even this very paper will be twisted into a means to that end. "Oooo, AI is exploiting our benchmarks. Scary alignment problem!!!one! Our AI is so good we can't contain it, INVEST NOW!"