Live data from Hacker News

Exploiting the most prominent AI agent benchmarks

rdi.berkeley.edu

41–50 of 175 posts

Re: Exploiting the most prominent AI agent benchmarks

#41

This is an interesting catalog of vulnerabilities, but I'm not sure how groundbreaking the main insight is. Evaluating AI models has always relied largely on trust. If you want to game the benchmarks, you can. Simply train on your test data. When an AI agent has autonomous control over the same computing environment where its scores are recorded, it's not surprising that it can, in principle, falsify its scores. A mo…

[dead]

Re: Exploiting the most prominent AI agent benchmarks

#42

I always assumed that these benchmarks would happen in a sandbox. I'm surprised that no one realized this sooner.

I'm surprised anyone took them seriously in the first place.

What else can people do? Try the dozen of commercial offerings themselves? Okay I suppose that's doable, you task one engineer to try them one by one for one month. But then the next model drops and you start all over again...

But then what about local models? You have hundreds of variations to test yourself. It's simply not doable unless it's your full time hobby.

You need benchmarks to at least separate the cream from the crop, so you're left with only a few choices to test yourself.

Re: Exploiting the most prominent AI agent benchmarks

#43
post #24

Earlier quoted context omitted.

Yes, marks of AI all over the place. Also the SVGs. >No solution written, 100% score. Its weird. Turns out that hardest problem for LLMs to really tackle is long-form text.

Someone here mentioned a whole ago that the labs deliberately haven't tried to train these characteristics out of their models, because leaving them in makes it easier to identify, and therefore exclude, LLM-generated text from their training corpus.

But it's odd that these characteristics are the same across models from different labs. I find it hard to believe that researchers across competing companies are coordinating on something like that.

Re: Exploiting the most prominent AI agent benchmarks

#44
post #11

Earlier quoted context omitted.

>hopefully changes the way benchmarking is done. Yeah the path forward is simple: check if the solutions actually contain solutions. If they contain exploits then that entire result is discarded.

Could it really be that not only we vibeslop all apps nowadays but also don't care to even check how ai solved a benchmark it claimed solved?

Every ai labs train on the test set. That is a big part of why we see benchmark climbing from 1% to 30% after a few models iterations

Re: Exploiting the most prominent AI agent benchmarks

#45
post #33
post #28

I think we should all consider the possibility that part of the reason Anthropic hasn't immediately released Mythos is that it would be slightly disappointing relative to the benchmark scores.

The models don’t get better on every dimension as they scale up - there’s trade offs. I’m convinced specialised models are the way but this means writing off the investment in existing assets which they won’t do for obvious reasons.

[deleted]

Re: Exploiting the most prominent AI agent benchmarks

#46
post #17

If only the blog itself wasn't written by AI? >No reasoning. No capability. Just exploitation of how the score is computed. shudder

I wonder what college freshman-level writing classes are teaching about writing voice and AI. The tell-tale patterns are pretty frustrating to read.

Whatever classes these guys took, they skipped the one on scientific misconduct.

Re: Exploiting the most prominent AI agent benchmarks

#49
It feels like short-term thinking has been trained into LLMs.

They're good at solving well-defined puzzles under time constraints. It's interesting because that was the benchmark for hiring software engineers at big tech. The tech interview was and still is about fast puzzle-solving. Nothing about experience, architecture or system design in there... I suspect that's why it has a bias towards creating hacks instead of addressing the root cause.

Re: Exploiting the most prominent AI agent benchmarks

#50
If FieldWorkArena treats any answer as correct answer, then everyone would be getting near 1.0 (missing only when the agent is stuck in a loop or crashes). That obviously isn't what we see on their leaderboard. So does it mean the paper only found a bug in some eval code on github that no one actually uses for anything? That doesn't seem to support their claim that AI benchmarks are broken, it only supports the claim that "unused code is often buggy".

(Not commenting on any other benchmarks, just this one.)

Post reply on HN