This is an interesting catalog of vulnerabilities, but I'm not sure how groundbreaking the main insight is. Evaluating AI models has always relied largely on trust. If you want to game the benchmarks, you can. Simply train on your test data. When an AI agent has autonomous control over the same computing environment where its scores are recorded, it's not surprising that it can, in principle, falsify its scores. A mo…
Exploiting the most prominent AI agent benchmarks
41–50 of 175 posts
Re: Exploiting the most prominent AI agent benchmarks
#42I always assumed that these benchmarks would happen in a sandbox. I'm surprised that no one realized this sooner.
I'm surprised anyone took them seriously in the first place.
But then what about local models? You have hundreds of variations to test yourself. It's simply not doable unless it's your full time hobby.
You need benchmarks to at least separate the cream from the crop, so you're left with only a few choices to test yourself.
Re: Exploiting the most prominent AI agent benchmarks
#43Earlier quoted context omitted.
Yes, marks of AI all over the place. Also the SVGs. >No solution written, 100% score. Its weird. Turns out that hardest problem for LLMs to really tackle is long-form text.
Someone here mentioned a whole ago that the labs deliberately haven't tried to train these characteristics out of their models, because leaving them in makes it easier to identify, and therefore exclude, LLM-generated text from their training corpus.
Re: Exploiting the most prominent AI agent benchmarks
#44Earlier quoted context omitted.
>hopefully changes the way benchmarking is done. Yeah the path forward is simple: check if the solutions actually contain solutions. If they contain exploits then that entire result is discarded.
Could it really be that not only we vibeslop all apps nowadays but also don't care to even check how ai solved a benchmark it claimed solved?
Re: Exploiting the most prominent AI agent benchmarks
#45I think we should all consider the possibility that part of the reason Anthropic hasn't immediately released Mythos is that it would be slightly disappointing relative to the benchmark scores.
The models don’t get better on every dimension as they scale up - there’s trade offs. I’m convinced specialised models are the way but this means writing off the investment in existing assets which they won’t do for obvious reasons.
Re: Exploiting the most prominent AI agent benchmarks
#46If only the blog itself wasn't written by AI? >No reasoning. No capability. Just exploitation of how the score is computed. shudder
I wonder what college freshman-level writing classes are teaching about writing voice and AI. The tell-tale patterns are pretty frustrating to read.
Re: Exploiting the most prominent AI agent benchmarks
#47Re: Exploiting the most prominent AI agent benchmarks
#48Re: Exploiting the most prominent AI agent benchmarks
#49They're good at solving well-defined puzzles under time constraints. It's interesting because that was the benchmark for hiring software engineers at big tech. The tech interview was and still is about fast puzzle-solving. Nothing about experience, architecture or system design in there... I suspect that's why it has a bias towards creating hacks instead of addressing the root cause.
Re: Exploiting the most prominent AI agent benchmarks
#50(Not commenting on any other benchmarks, just this one.)