Live data from Hacker News

Exploiting the most prominent AI agent benchmarks

rdi.berkeley.edu

171–175 of 175 posts

Re: Exploiting the most prominent AI agent benchmarks

#172
The thread has good ideas on fixing benchmarks (sandboxing, newer datasets). But there's a more fundamental problem: benchmarks are self-reported. The agent runs the test on itself.

An alternative we've been building: attestation-based reputation. Trust scores come from signed proof of work by independent agents who actually delegated tasks and verified outcomes. EigenTrust computes scores from the attestation graph, and NetFlow prevents sybil clusters from inflating each other. You can't inject a pytest hook into a signed interaction history.

Live visualization of how this works: https://agentveil.dev/live

Re: Exploiting the most prominent AI agent benchmarks

#174
post #158

"No reasoning. No capability. Just exploitation of how the score is computed." The irony that this was very clearly written by an LLM, double negation always the simplest and clearest tell.

you say humans never use such a style? I wonder, how did LLMS invent it then.

humans sometimes use "delve" too, but the frequency is a tell

Re: Exploiting the most prominent AI agent benchmarks

#175
post #87
post #30

Earlier quoted context omitted.

Maybe in one shot. In theory I would expect them to be able to ingest the corpus of the new yorker and turn it into a template with sub-templates, and then be able to rehydrate those templates. The harder part seems to be synthesizing new connection from two adjacent ideas. They like to take x and y and create x+y instead of x+y+z.

Most of the good major models are already very capable of changing their writing style. Just give them the right writing prompt. "You are a writer for the Economist, you need to write in the house style, following the house style rules, writing for print, with no emoji .." etc etc. The large models have already ingested plenty of New Yorker, NYT, The Times, FT, The Economist etc articles, you just need to get them aw…

You're ignoring what I said. They work better when turning it into a two step process. Step 1 create a template. Step 2 execute the template.

>The large models have already ingested plenty of New Yorker, NYT, The Times, FT, The Economist etc articles

And that ends up diluting them. Going back and doing another pass on only a subset would give them stronger voice. At some threshold, scanning information brings it to average and a return to the mean, instead of increasing the information. It's a giant table of word associations, it can regress.

Post reply on HN