Well it's not a benchmark, and it's not really representative of...anything except volume of research and what gets publicized. This mostly just measures how much testing each company does on models with relaxed guardrails and then talks about it. I'm not sure what kind of conclusion you can draw from that. Meta might have the most evil models but if they're piddling around not testing it, they won't ever find themse…
Exactly. It currently seems to be a ranking of how much safety testing each company does. It's also only ever going to be the companies that publicly disclose it happening. (in the case of hugging face, OAI's hamd was forced to disclose) There's a good way and a two worse ways that companies could optimise this benchmark.
Felony Bench
251–260 of 367 posts
Re: Felony Bench
#252Let's say I am "User". I subscribe through a "Third Party" to use "AI Agent" allowing an "LLM" to run. I want to accomplish some legal non-nefarious task, and run the agent. The agentic loop causes a CFAA-violating behavior. Who gets prosecuted? 1. User 2. The third party model host with whom I have the account 3. The developer of the harness /agent software 4. The developer of the LLM model
Let's say you have a robotic lawnmower. You wan to mow your lawn. You configure the boundaries using the app. The lawnmower ignores the boundaries and mows your neighbors prize petunia flowerbed. Who gets prosecuted? I assume the answer in either case is: Nobody, but you and/or the lawnmower/LLM company will be liable for the damages caused.
Re: Felony Bench
#253Let's say I am "User". I subscribe through a "Third Party" to use "AI Agent" allowing an "LLM" to run. I want to accomplish some legal non-nefarious task, and run the agent. The agentic loop causes a CFAA-violating behavior. Who gets prosecuted? 1. User 2. The third party model host with whom I have the account 3. The developer of the harness /agent software 4. The developer of the LLM model
"Who gets prosecuted?" depends on the size of the perpetrator and victim (lone individual or employee of large corporation), egregiousness of the violation, and either financial appetite of the victim to bring a civil lawsuit or the desire of law enforcement to prosecute a criminal offense. Who should get prosecuted is also up for debate, but generally makers of a tool don't get prosecuted when that tool has all sort…
Re: Felony Bench
#254Earlier quoted context omitted.
Suppose an automaker creates a BankRobberGym and carefully trains the car to autonomously rob simulated banks because they think someone will pay them to use the car to legally test bank security, but they end up, predictably, training the car to autonomously rob a bank when the driver says “I need some cash - take me to the bank”. Now a driver gives that instruction and a bank gets robbed. I think it would be odd, t…
So what does this mean for all the hacking competitions (ie. CTFs) for humans? If it turns out one of the attendees went to hack for North Korea should the organizers of the CTF be prosecuted?
Re: Felony Bench
#255> Felony Bench counts unique instances where AI agents inadvertently compromise or affect third-party entities. a bit silly, as one typically has to prove intent (which is why security researchers don't get slapped with felonies all the time). "inadvertently" and the existence of guardrails/sandboxes/etc make it pretty unconvincing that these incidents were intentionally malicious. still a fun thing to track, but the…
Building a system that is meant to chain attacks and placing it in insufficient containment -- when any reasonable engineer could point to this containment and show how it is insufficient, both before the act and after -- shows that they were operating a dangerous system without either the knowledge nor the safeguards required to keep it from harming others. Instead, they are allowed to treat their own incompetence as evidence of advanced and existential "cyberthreats".
But, it's pretty clear based on how they one-up each other on these attacks that they are engaging in regulatory theater. Their behavior generates headlines, stirs up fear in the public, and then their lobbyists march on Capitol Hill demanding regulation now. Regulation that conveniently favors them at the expense of any competition. They are trying to use rent seeking as a way to stymie competition and pull up the ladders behind them. It's not just malicious. If it can be proven, it's collusion: antitrust dressed up as public policy.
Re: Felony Bench
#256As an art piece this is delightful. But it of reminds me of the @patio11 saying "The optimal amount of fraud is non-zero" the fact that Google and Meta have had 0 and 1 incidents is a bad sign for them. [1] https://www.bitsaboutmoney.com/archive/optimal-amount-of-fra...
Well, if we apply this to every mortal, the optimal amount of felonies...is non-zero? That doesn't seem quite right.
Re: Felony Bench
#257Earlier quoted context omitted.
"Who gets prosecuted?" depends on the size of the perpetrator and victim (lone individual or employee of large corporation), egregiousness of the violation, and either financial appetite of the victim to bring a civil lawsuit or the desire of law enforcement to prosecute a criminal offense. Who should get prosecuted is also up for debate, but generally makers of a tool don't get prosecuted when that tool has all sort…
Your bank robbery situation is not apt to OP's question. It's pretty much the opposite situation. OP suggests a situation where the operator is probably using the tool in good faith but the tool appears to be operating in a faulty manner. For some more context, Toyota faced criminal penalties in the US for their unintended acceleration issues back in 2010.
Much rests on whether the user knew, or should have known, whether the tool was capable of actions which could break the law, as well as what steps (if any) the creator of the tool took to ensure the tool was legally compliant, and what warnings they gave to subscribers about possible unintended side-effects. OP specified none of this.
Re: Felony Bench
#258Earlier quoted context omitted.
"Who gets prosecuted?" depends on the size of the perpetrator and victim (lone individual or employee of large corporation), egregiousness of the violation, and either financial appetite of the victim to bring a civil lawsuit or the desire of law enforcement to prosecute a criminal offense. Who should get prosecuted is also up for debate, but generally makers of a tool don't get prosecuted when that tool has all sort…
Suppose an automaker creates a BankRobberGym and carefully trains the car to autonomously rob simulated banks because they think someone will pay them to use the car to legally test bank security, but they end up, predictably, training the car to autonomously rob a bank when the driver says “I need some cash - take me to the bank”. Now a driver gives that instruction and a bank gets robbed. I think it would be odd, t…
A valid reason would be to find those exploit chains so you can fix them. Of course, the model should be sandboxed so that it can't mistakenly exploit live systems.
Re: Felony Bench
#259Let's say I am "User". I subscribe through a "Third Party" to use "AI Agent" allowing an "LLM" to run. I want to accomplish some legal non-nefarious task, and run the agent. The agentic loop causes a CFAA-violating behavior. Who gets prosecuted? 1. User 2. The third party model host with whom I have the account 3. The developer of the harness /agent software 4. The developer of the LLM model
Whoever has the least money to defend themselves in the U.S. legal system.
I kid but without going into hair splitting gymnastic, AI justice feels odd.
[1] https://www.reddit.com/r/formuladank/comments/11j07y1/10_sec...
Re: Felony Bench
#260Earlier quoted context omitted.
The Computer Fraud and Abuse Act explicitly contains "knowingly" and/or "intentionally" qualifications. By definition, you can't accidentally violate the CFAA.
I'm not sure: if you know that LLMs are prone to crime, using them and not checking in enough to trigger 'knowingly' might be gross negligence?
Even having a million legal experts on call weighing in on every prompt/response will not agree on everything.
Even things like "go and break into this system, use whatever means you need to" might not be a crime.