Earlier quoted context omitted.
I get your point, but it’s obvious that everybody is going to do this. So building a benchmark isn’t likely to directly move capabilities. OTOH it might give some good signals on required changes in training recipes.
People are still debating the liabilities associated with AI development. If your agent hacks the CIA, people want to blame the AI lab. but... If your agent spends $100k on tokens, then thats a user error. If your agent spends $10m on a shopping spree, then that is also a user error?
At some point these systems will get certified as fiduciary agents but they sure as hell aren’t claimed to be that now.