Live data from Hacker News

GPT-5 outperforms federal judges in legal reasoning experiment

papers.ssrn.com

91–100 of 254 posts

Re: GPT-5 outperforms federal judges in legal reasoning experiment

#92

Terrifying concept this is literally saying if AI was legal we'd have an absolute rigid dystopia

And this was just about how to decide an auto accident case. With the experiment varying the circumstances.

My summary is still: seasoned judges disagree with LLM output 50% of the time.

Re: GPT-5 outperforms federal judges in legal reasoning experiment

#93
post #87

IANAL, but this seems like an odd test to me. Judges do what their name implies - make judgment calls. I find it re-assuring that judges get different answers under different scenarios, because it means they are listening and making judgment calls. If LLMs give only one answer, no matter what nuances are at play, that sounds like they are failing to judge and instead are diminishing the thought process down to black-…

> Judges do what their name implies - make judgment calls. I find it re-assuring that judges get different answers under different scenarios, because it means they are listening and making judgment calls. I disagree - law should be the same for everyone. Yes sometimes crimes have mitigating curcumstances and those should be taken into account. However that seems like a separate question of what is and is not illegal.

The thing is, Laws do not forsee in all cases, and language is not completely objective, so you cannot avoid judgement calls. One example is computer hacking, which in many jurisdictions is specified in very vague terms.

Re: GPT-5 outperforms federal judges in legal reasoning experiment

#94

IANAL, but this seems like an odd test to me. Judges do what their name implies - make judgment calls. I find it re-assuring that judges get different answers under different scenarios, because it means they are listening and making judgment calls. If LLMs give only one answer, no matter what nuances are at play, that sounds like they are failing to judge and instead are diminishing the thought process down to black-…

I don't think a lot of people understand the grueling nature of a judge. Day in and out of cases over years are going to generate bias in the judge in one form or another. I wouldn't mind an AI check* to help them check that bias

*A magically thorough, secure, and well tested AI

Re: GPT-5 outperforms federal judges in legal reasoning experiment

#95
post #71

Earlier quoted context omitted.

Yeah, I'm reminded of the various child porn cases where the "perpetrator" is a stupid teenager who took nude pics of themselves and sent them to their boy/girlfriend. Many of those cases have been struck down by judges because the letter of the law creates a non-sequitur where the teenager is somehow a felon child predator who solely preyed on themselves, and sending them to jail and forcing them to sign up for a se…

This is one of the roles of justice, but it is also one of the reasons why wealthy people are convicted less often. While it often delivered as a narrative of wealth corrupting the system, the reality is that usually what they are buying is the justice that we all should have. So yes, a judge can let a stupid teenager off on charges of child porn selfies. but without the resources, they are more likely be told by a p…

Sure, but I'm not sure how AI would solve any of that.

Any claims of objectivity would be challenged based on how it was trained. Public opinion would confirm its priors as it already does (see accusations of corruption or activism with any judicial decision the mob disagrees with, regardless of any veracity). If there's a human appeals process above it, you've just added an extra layer that doesn't remove the human corruption factor at all.

As for corruption, in my opinion we're reading some right now. Human-in-the-loop AI doesn't have the exponential, world-altering gains that companies like OpenAI need to justify their existence. You only get that if you replace humans completely, which is why they're all shilling science fiction nonsense narratives about nobody having to work. The abstract of this paper leans heavily into that narrative

Re: GPT-5 outperforms federal judges in legal reasoning experiment

#96
post #87

IANAL, but this seems like an odd test to me. Judges do what their name implies - make judgment calls. I find it re-assuring that judges get different answers under different scenarios, because it means they are listening and making judgment calls. If LLMs give only one answer, no matter what nuances are at play, that sounds like they are failing to judge and instead are diminishing the thought process down to black-…

> Judges do what their name implies - make judgment calls. I find it re-assuring that judges get different answers under different scenarios, because it means they are listening and making judgment calls. I disagree - law should be the same for everyone. Yes sometimes crimes have mitigating curcumstances and those should be taken into account. However that seems like a separate question of what is and is not illegal.

The law is rife with words and phrasing that make legality dependent upon those subjective mitigating factors.

Re: GPT-5 outperforms federal judges in legal reasoning experiment

#97
post #87

IANAL, but this seems like an odd test to me. Judges do what their name implies - make judgment calls. I find it re-assuring that judges get different answers under different scenarios, because it means they are listening and making judgment calls. If LLMs give only one answer, no matter what nuances are at play, that sounds like they are failing to judge and instead are diminishing the thought process down to black-…

> Judges do what their name implies - make judgment calls. I find it re-assuring that judges get different answers under different scenarios, because it means they are listening and making judgment calls. I disagree - law should be the same for everyone. Yes sometimes crimes have mitigating curcumstances and those should be taken into account. However that seems like a separate question of what is and is not illegal.

Laws are written to be interpreted and applied by humans. They aren’t computer programs. They are full of ambiguity. Much of this is by design because there are too many possible edge cases to design a fully algorithmic unambiguous legal system.

Re: GPT-5 outperforms federal judges in legal reasoning experiment

#98
post #23

IANAL, but this seems like an odd test to me. Judges do what their name implies - make judgment calls. I find it re-assuring that judges get different answers under different scenarios, because it means they are listening and making judgment calls. If LLMs give only one answer, no matter what nuances are at play, that sounds like they are failing to judge and instead are diminishing the thought process down to black-…

> If LLMs give only one answer, no matter what nuances are at play, that sounds like they are failing to judge and instead are diminishing the thought process down to black-and-white thinking. You can have a team of agents exchange views and maybe the protocol would even allow for settling the cases automatically. The more agents you have, the higher the nuances.

Then you'd need to provide them with access to the law, previous cases, to the news, to various data sources. And you'd have to decide how much each of those sources of information matter. And at that point, you've got people making the decision again instead of the ai in practice.

And then there's the question of the model used. Turns out I've got preferences for which model I'd rather be judged by, and it's not Grok for example...

Re: GPT-5 outperforms federal judges in legal reasoning experiment

#99

The premise seems flawed. From the paper: “we find that the LLM adheres to the legally correct outcome significantly more often than human judges” That presupposes that a “legally correct” outcome exists The Common Law, which is the foundation of federal law and the law of 49/50 states, is a “bottom up” legal system. Legal principals flow from the specific to the general. That is, judges decided specific cases based…

Arguing that this is a Common Law matter in this scenario is funny in a wonky lawyerly kind of way.

The legal issue they were testing in this experiment is choice of law and procedure question, which is governed by a line of cases starting with Erie Railroad in which Justice Brandies famously said, "There is no federal common law."

Re: GPT-5 outperforms federal judges in legal reasoning experiment

#100
post #37

Earlier quoted context omitted.

The legal system leaves much to be desired in relation to fairness and equity. I’d much prefer a multi-staged approach with an 1) AI analysis, 2) judge review with high bar for analysis if in disagreement with the AI, 3) public availability of the deliberations, 4) an appeals process.

Even having a ready-made determination by an AI runs the risk of prejudicing judges and juries.

Given TFA, it seems that having human determinations involved might run the risk of prejudicing the AI.
Post reply on HN