Live data from Hacker News

Felony Bench

felonybench.com

271–280 of 367 posts

Re: Felony Bench

#271

Earlier quoted context omitted.

The Computer Fraud and Abuse Act explicitly contains "knowingly" and/or "intentionally" qualifications. By definition, you can't accidentally violate the CFAA.

OpenAI and Anthropic both have currently safety teams that look for misbehavior in their models (and to some extent, voluntarily disclose what they find to the public). Going forward, it would be hard for them to argue they don’t know their models do stuff like this.

But you could also use this to argue in the other way to say that they are using due care and therefore not negligent

Re: Felony Bench

#272
post #249

Earlier quoted context omitted.

So what does this mean for all the hacking competitions (ie. CTFs) for humans? If it turns out one of the attendees went to hack for North Korea should the organizers of the CTF be prosecuted?

There’s a difference between enabling another individual with free will and agency, and enabling an automated tool (as a bonus, then giving it to the masses & profiting from its use).

There are certainly some differences but is one really more ethical than the other? If anything I feel that (for example) manufacturing a gun is much less likely to carry any ethical implications than training someone to use it might.

Re: Felony Bench

#273
post #267
post #92

Let's say I am "User". I subscribe through a "Third Party" to use "AI Agent" allowing an "LLM" to run. I want to accomplish some legal non-nefarious task, and run the agent. The agentic loop causes a CFAA-violating behavior. Who gets prosecuted? 1. User 2. The third party model host with whom I have the account 3. The developer of the harness /agent software 4. The developer of the LLM model

Cause-and-effect could quickly turn into butterfly effect. Let's say you were fixing a screw on a device in a low light conditions, the screw head is badly manufactured and the screwdriver isn't made according to standards, the tool breaks and flies away, bounces off a bench which shouldn't be there and hits someone who is roaming in the workplace unauthorized and without following safety rules. Now, who do you blame…

Sounds like you'll need to give all your money to a team of lawyers and wait a few years to get an answer. /s

But for LLM stuff most non-contrived examples are actually fairly trivial. Try replacing "LLM" with "self driving car" and see if that helps. Basically ask was the operator negligent, was a bystander negligent, were the vendor or manufacturer negligent, etc.

Re: Felony Bench

#274
post #232
post #123

Earlier quoted context omitted.

"Who gets prosecuted?" depends on the size of the perpetrator and victim (lone individual or employee of large corporation), egregiousness of the violation, and either financial appetite of the victim to bring a civil lawsuit or the desire of law enforcement to prosecute a criminal offense. Who should get prosecuted is also up for debate, but generally makers of a tool don't get prosecuted when that tool has all sort…

Suppose an automaker creates a BankRobberGym and carefully trains the car to autonomously rob simulated banks because they think someone will pay them to use the car to legally test bank security, but they end up, predictably, training the car to autonomously rob a bank when the driver says “I need some cash - take me to the bank”. Now a driver gives that instruction and a bank gets robbed. I think it would be odd, t…

> In regard to “cyber”, there is, IMO, no valid reason whatsoever to train a model to autonomously create exploit chains.

Field testing is a real thing in literally all industries.

Except, apparently, the software industry. When it comes to software security and protecting your sensitive data, the solution is "trust me bro, I got my team of the best lawyers on it".

Re: Felony Bench

#275

These are just cases of AI models committing illegal activity - without any legal convictions yet. If that's the logic, how is Grok not at the top of the list for deepfaking millions? Edit: I get that this is about agents, but a lot of these instances are about agents going rogue after the human gave them a task. "inadvertently" breaking the law isn't necessarily a lesser category than "did so on command." If we are…

[dead]

Re: Felony Bench

#277

Earlier quoted context omitted.

charged is possible, however i doubt there would be a conviction for the reasons i stated (no intent).

I think it’s most likely you’re right, but I’m the weirdo who thinks there’s actually a non-zero probability that there was negative intent and that the “accidental” aspect is a form of damage control. If someone broke into my house but then claimed they didn’t mean to when they saw I was home, I’m not sure I’d take them at their word.

You think that a training run in a VM that was set up with access only to an internal package repository was intended to hack said repository to then go on and hack HF?

All so that OpenAI could do a bit of bragging and massively delay their own work and planned model rollouts?

That seems extremely unlikely.

Re: Felony Bench

#278
post #43

The way that OpenAI has communicated around the HuggingFace incident makes me feel crazy. You created a machine that undertook a malicious campaign of harm against an innocent third-party! You should be doing deep introspection about how your company culture and approach to R&D produces criminal outcomes. Instead, they treat their own felonious behavior like it is an uncontrollable act of God. From Greg Brockman's po…

OpenAI and Anthropic are falling over themselves to claim these "incidents" show their products are both amazingly super-powerful and also "dangerous" so they need to be regulated. In addition to these stories, these companies are sponsoring "please regulate us" ads. ( https://www.cnbc.com/2026/02/19/dueling-pacs-take-center-sta... ) Like Uber, companies that had no concern for the law as they innovated their way to the top, once there, push for laws to limit competition.

Re: Felony Bench

#279

These are just cases of AI models committing illegal activity - without any legal convictions yet. If that's the logic, how is Grok not at the top of the list for deepfaking millions? Edit: I get that this is about agents, but a lot of these instances are about agents going rogue after the human gave them a task. "inadvertently" breaking the law isn't necessarily a lesser category than "did so on command." If we are…

I don't think this "benchmark" is about alignment, per se.

I think it's more about: presuming alignment failure happens, then how many exploits will each given model implicitly come up with and use; how many systems will it implicitly break out of and through and into; and how many laws will it implicitly end up violating, all in the process of trying to accomplish some non-aligned sub-goal (e.g. "cheating" at its answer) of the prompt you've given it, all during a single conversation turn, without asking for any additional user input or confirmations?

In other words, how big a rocket-powered sledgehammer does the model have sitting around in its golf bag, just waiting for it to decide to give it a swing the next time you attempt to swat a fly?

Post reply on HN