Exploiting the most prominent AI agent benchmarks
171–175 of 175 posts
Re: Exploiting the most prominent AI agent benchmarks
#172An alternative we've been building: attestation-based reputation. Trust scores come from signed proof of work by independent agents who actually delegated tasks and verified outcomes. EigenTrust computes scores from the attestation graph, and NetFlow prevents sybil clusters from inflating each other. You can't inject a pytest hook into a signed interaction history.
Live visualization of how this works: https://agentveil.dev/live
Re: Exploiting the most prominent AI agent benchmarks
#173Re: Exploiting the most prominent AI agent benchmarks
#174"No reasoning. No capability. Just exploitation of how the score is computed." The irony that this was very clearly written by an LLM, double negation always the simplest and clearest tell.
you say humans never use such a style? I wonder, how did LLMS invent it then.
Re: Exploiting the most prominent AI agent benchmarks
#175Earlier quoted context omitted.
Maybe in one shot. In theory I would expect them to be able to ingest the corpus of the new yorker and turn it into a template with sub-templates, and then be able to rehydrate those templates. The harder part seems to be synthesizing new connection from two adjacent ideas. They like to take x and y and create x+y instead of x+y+z.
Most of the good major models are already very capable of changing their writing style. Just give them the right writing prompt. "You are a writer for the Economist, you need to write in the house style, following the house style rules, writing for print, with no emoji .." etc etc. The large models have already ingested plenty of New Yorker, NYT, The Times, FT, The Economist etc articles, you just need to get them aw…
>The large models have already ingested plenty of New Yorker, NYT, The Times, FT, The Economist etc articles
And that ends up diluting them. Going back and doing another pass on only a subset would give them stronger voice. At some threshold, scanning information brings it to average and a return to the mean, instead of increasing the information. It's a giant table of word associations, it can regress.