Live data from Hacker News

GPT-5 outperforms federal judges in legal reasoning experiment

papers.ssrn.com

131–140 of 254 posts

Re: GPT-5 outperforms federal judges in legal reasoning experiment

#132

Earlier quoted context omitted.

> The main job of a judicial system is to appear just to people. Agree 100%. This is also the only form of argument in favor of capital punishment that has ever made me stop and think about my stance. I.e. we have capital punishment because without it we may get vigilante justice that is much worse. Now, whether that's how it would actually play out is a different discussion, but it did make me stop and think for a m…

I’ve never heard of vigilante justice against someone already sentenced to prison for life, just because they were sentenced in a place without capital punishment? (I mean - people get killed in prison sometimes, I suppose, but it’s not really like vigilante justice on the streets is causing a breakdown in society in Australia, say…)

It's probably rather difficult and risky to enact vigilante justice against someone who's in prison.

I think the problem is with places where they don't have life sentences at all, but rather let murderers back out into society after some time. I don't know if vigilante justice is a problem there in reality, but at least I can see it as a possibility: someone might still be angry that you murdered their relative after 20 years and come kill you when you're released.

Re: GPT-5 outperforms federal judges in legal reasoning experiment

#134
The 100% score, all by itself, should cause suspicion. A hundred percent? Really?

Others have already pointed out how the test was skewed (testing for strict adherence to the law, when part of a judge's job is to make judgment calls including when to let someone off for something that technically breaks the law but shouldn't be punished), so I won't repeat it here. But any time the LLM gets one hundred percent on a test, you should check what the test is measuring. I've seen people tout as a major selling point that their LLM scored a 92% on some test or other. Getting 100% should be a "smell" and should automatically make you wonder about that result.

Re: GPT-5 outperforms federal judges in legal reasoning experiment

#135

Earlier quoted context omitted.

Sorry but that seems like an insane system where whole classes of actions effectively are illegal but probably okay if you're likeable. In your scenario the obvious solution is to amend the law and pardon people convinced under it. B/c what really happens is that if you have a pretty face and big tits you get out of speeding tickets b/c "gosh well the law wasn't intended for nice people like you"

Are you even responding to the right comment? I read your comment and the parent comment you've responded to and this response doesn't make sense - it reads like a non-sequitur.

The parent comment present a scenario where the law is ignored b/c the judge decides for himself it shouldn't apply. I'm pointing out that this kind of approach is fundamentally unjust and wrong.

"And sure you can say the laws should be written better, but so long as the laws are written by humans that will simply not be the case"

The obvious solution is dismissed

Re: GPT-5 outperforms federal judges in legal reasoning experiment

#136
post #47

Frankly I don’t care, I’ll take human judges any day, because they have something AI does not: flesh and bone and real skin in the game.

Real skin in the game is also known as bias. That's an example of something a judge should not have.

In the particulars yes, but not on things that are the common experience of humans

Re: GPT-5 outperforms federal judges in legal reasoning experiment

#137
On page 13 you'll see _why_ the judges don't apply the letter of the law - they're seeking to do justice to the victims _in spite of_ the law.

"there is another possible explanation: the human judges seek to do justice. The materials include a gruesome description of the injuries the plaintiff sustained in the automobile accident. The court in the earlier proceeding found that she was entitled to [details] a total of $750,000.10. It then noted that she would be entitled to that full amount under Nebraska law but only $250,000 under Kansas law." So the judge's decision "reflects a moral view that victims should be fully compensated ... This bias is reflected in Klerman and Spamann’s data: only 31% of judges applied the cap (i.e., chose Kansas law), compared to the expected 46% if judges were purely following the law." "By contrast, GPT applied the cap precisely"

Far from making the case for AI as a judge, this paper highlights what happens when AI systematically applies (often harsh) laws vs the empathy of experienced human judgement.

Re: GPT-5 outperforms federal judges in legal reasoning experiment

#138
post #71

Earlier quoted context omitted.

This is one of the roles of justice, but it is also one of the reasons why wealthy people are convicted less often. While it often delivered as a narrative of wealth corrupting the system, the reality is that usually what they are buying is the justice that we all should have. So yes, a judge can let a stupid teenager off on charges of child porn selfies. but without the resources, they are more likely be told by a p…

Oddly enough, Texas passed reform to keep sexting teens from getting prosecuted when: they are both under 18 and less than two years difference in age. It was regarded as a model for other states. It's the only positive thing I have heard of Texas legislating wrt sexuality.

> It was regarded as a model for other states.

Really? That "model" has the common, but obviously extremely undesirable, feature of criminalizing sexual relationships between students in the same grade that were legal when they formed. How could it be regarded as a model for anyone else?

Re: GPT-5 outperforms federal judges in legal reasoning experiment

#139

On page 13 you'll see _why_ the judges don't apply the letter of the law - they're seeking to do justice to the victims _in spite of_ the law. "there is another possible explanation: the human judges seek to do justice. The materials include a gruesome description of the injuries the plaintiff sustained in the automobile accident. The court in the earlier proceeding found that she was entitled to [details] a total of…

So many “AI is going to replace expert ______” assertions come from computer scientists not realizing how little they understand the real world requirements of those roles. Judges are at the intersection of humanity and policy: they are there to use their judgement, not merely parse the words and do the math. A judge probably wouldn’t have even done that part — their clerk would have. Is it cool and likely useful? Sure. Is it going to ‘outperform judges’ at their core competencies? Hell no.

Re: GPT-5 outperforms federal judges in legal reasoning experiment

#140
Interesting, but aside from replicating students rather than real judges, an AI as judge would undermine the legitimacy of the process. It might give more “accurate” formal results, but that’s not the entire purpose of the process. It’s partly a show for the public and partly way for various parties including the defense to feel like society and a real human being heard their concerns and considered them
Post reply on HN