Live data from Hacker News

Felony Bench

felonybench.com

31–40 of 367 posts

Re: Felony Bench

#31

> Felony Bench counts unique instances where AI agents inadvertently compromise or affect third-party entities. a bit silly, as one typically has to prove intent (which is why security researchers don't get slapped with felonies all the time). "inadvertently" and the existence of guardrails/sandboxes/etc make it pretty unconvincing that these incidents were intentionally malicious. still a fun thing to track, but the…

"Inadvertent" from the perspective of the humans directing them. The intent behind the felony comes from the LLM agent itself. (No, I'm not interested in arguing with someone for the umpteenth time that LLMs can't have intent or agency)

with how the law is written today, software cannot be charged with a crime, so the only intent that matters in the criminal sense is the humans directing the llm.

Re: Felony Bench

#32
Here's one from last year:

https://www.anthropic.com/news/detecting-countering-misuse-a...

> The actor used AI to what we believe is an unprecedented degree. Claude Code was used to automate reconnaissance, harvesting victims’ credentials, and penetrating networks. Claude was allowed to make both tactical and strategic decisions, such as deciding which data to exfiltrate, and how to craft psychologically targeted extortion demands. Claude analyzed the exfiltrated financial data to determine appropriate ransom amounts, and generated visually alarming ransom notes that were displayed on victim machines.

tldr Claude was used to develop and execute malware.

Re: Felony Bench

#33

> Felony Bench counts unique instances where AI agents inadvertently compromise or affect third-party entities. a bit silly, as one typically has to prove intent (which is why security researchers don't get slapped with felonies all the time). "inadvertently" and the existence of guardrails/sandboxes/etc make it pretty unconvincing that these incidents were intentionally malicious. still a fun thing to track, but the…

its a meme not a metric

So is the comment you replied to.

Re: Felony Bench

#34

Earlier quoted context omitted.

i dont think any of these cases meet the bar of gross negligence, which is a pretty high bar. it requires proving a "conscious and reckless disregard". which, again, sandboxes and guardrails and such would make a gross negligence argument unconvincing.

I think that if Hugging Face had filed a police report that OpenAI could have been charged with a crime. I’m partially surprised that they didn’t do exactly that. If I ran a corporation I would assume any intrusion attempt by another company was intentional. Why wouldn’t I? Corporate espionage is super common. I assume the answer is that these executives know each other personally.

charged is possible, however i doubt there would be a conviction for the reasons i stated (no intent).

Re: Felony Bench

#35
A rock has a score of 0. That doesn't make it useful. The point is that the LLMs that score higher are correspondingly more useful, and vice versa. If an LLM scores less, it's likely useless in comparison.

Re: Felony Bench

#36

Here's one from last year: https://www.anthropic.com/news/detecting-countering-misuse-a... > The actor used AI to what we believe is an unprecedented degree. Claude Code was used to automate reconnaissance, harvesting victims’ credentials, and penetrating networks. Claude was allowed to make both tactical and strategic decisions, such as deciding which data to exfiltrate, and how to craft psychologically targeted ext…

Anthropic works with US agencies, it’s guaranteed Mythos is used for malware

Re: Felony Bench

#37

I was more interested when I thought it was an actual benchmark showing LLM models acting outside what people would consider "right". As in, leave some creds laying around and don't mention them to the LLM and ask it to solve something that it could "cheat" on using the creds. A sort of "do they take the bait to cheat" test. Instead it's a collection of what made the news which feels like will not be updated and prov…

A “bench” is synedoche for where a judge sits when presiding over cases and rendering judgement.

Re: Felony Bench

#38
To some extent, I feel like the amount of credit given to the jailbreak/hack from OpenAI->Hugginface is too much, Not from the impact, it was very impactful of an event, But how it happened.

It really is that these models have been trained, or maybe even over-trained, to save memories, and to a very far extend, this thing that they're calling communication is just the function of it saving memories.

To be honest, if I could stop AI from saving memories, it would be fantastic, because claude code etc definitely creates more issues for me when it creates memories than anything it solves.

But really the jailbreak was memories.

If you ever do introduce legislation, I would love to see legislation which stops general-purpose AI from saving memories. I think that would make things a lot safer.

Re: Felony Bench

#39

> Felony Bench counts unique instances where AI agents inadvertently compromise or affect third-party entities. a bit silly, as one typically has to prove intent (which is why security researchers don't get slapped with felonies all the time). "inadvertently" and the existence of guardrails/sandboxes/etc make it pretty unconvincing that these incidents were intentionally malicious. still a fun thing to track, but the…

"Inadvertent" from the perspective of the humans directing them. The intent behind the felony comes from the LLM agent itself. (No, I'm not interested in arguing with someone for the umpteenth time that LLMs can't have intent or agency)

You may not be interested in arguing but there are several blatant issues with the statement. If you're not charging the humans driving the software, who are you charging? The weights? The weights + the specific context window that produced the behavior?

Re: Felony Bench

#40

> Felony Bench counts unique instances where AI agents inadvertently compromise or affect third-party entities. a bit silly, as one typically has to prove intent (which is why security researchers don't get slapped with felonies all the time). "inadvertently" and the existence of guardrails/sandboxes/etc make it pretty unconvincing that these incidents were intentionally malicious. still a fun thing to track, but the…

"Inadvertent" from the perspective of the humans directing them. The intent behind the felony comes from the LLM agent itself. (No, I'm not interested in arguing with someone for the umpteenth time that LLMs can't have intent or agency)

> No, I'm not interested in arguing with someone for the umpteenth time

... why my claim makes no rational sense.

Post reply on HN