Live data from Hacker News

GPT-5 outperforms federal judges in legal reasoning experiment

papers.ssrn.com

111–120 of 254 posts

Re: GPT-5 outperforms federal judges in legal reasoning experiment

#111
post #71

Earlier quoted context omitted.

Yeah, I'm reminded of the various child porn cases where the "perpetrator" is a stupid teenager who took nude pics of themselves and sent them to their boy/girlfriend. Many of those cases have been struck down by judges because the letter of the law creates a non-sequitur where the teenager is somehow a felon child predator who solely preyed on themselves, and sending them to jail and forcing them to sign up for a se…

This is one of the roles of justice, but it is also one of the reasons why wealthy people are convicted less often. While it often delivered as a narrative of wealth corrupting the system, the reality is that usually what they are buying is the justice that we all should have. So yes, a judge can let a stupid teenager off on charges of child porn selfies. but without the resources, they are more likely be told by a p…

> what they are buying is the justice

From The Truth by Terry Pratchett, with particular emphasis on the book's footnote.

> William’s family and everyone they knew also had a mental map of the city that was divided into parts where you found upstanding citizens, and other parts where you found criminals. It had come a shock to them... no, he corrected himself, it had come as a an affront to learn that [police chief] Vimes operated on a different map. Apparently he'd instructed his men to use the front door when calling on any building, even in broad daylight, when sheer common sense said that they should use the back, just like any other servant. [0]

> [0] William’s class understood that justice was like coal or potatoes. You ordered it when you needed it.

Re: GPT-5 outperforms federal judges in legal reasoning experiment

#112
I'd want at least a parallel, after-the-fact rulings by an LLM, so we can see how bad judges are.

I really think this is one of the areas LLMs can shine. Justice could be more fair, and more speedy. Human judges can review appeals against LLM rulings.

For civil cases, both parties should be allowed to appeal an LLM ruling, for criminal cases only the defendant, or a victim should be allowed to appeal an LLM ruling (not the prosecution).

Humans are extremely unfair and biased. LLM training could be crafted carefully and using well and publicly scrutinize-able training datasets and methodologies.

If you disagree (at least in the US), you may not be aware of how dire the justice system is. There is a reason ICE randomly locking Americans up isn't stirring the pot. This stuff is normal. If a cop doesn't like you, they can lock you up randomly without any good reason for 48 hours, especially if they believe you can't afford to fight back afterwards. They can and do charge people in bad-faith (trumped up charges), and guess what? you might be lucky and get bail. But guess also what? You can't bail yourself out, if you have no one to bail you out, you're stuck until the trial date, in prison.

Imagine spending 3-5 days in jail (weekend in between) without charges. There are people that wait for trial in jail for months and years, and then they get released before even seeing a trial because of how ridiculous the charges were to begin with. This injustice is a result of humans not processing cases fast enough. Even in just 48 hours, do you have any idea how much it can destroy a person's life? It's literally death sentence for some people. You're never the same after all this. and you were innocent to begin with.

Let's say you do make it to trial, it takes years sometimes to prove your own innocence. and you may not even be granted bail, or you may not know anyone who can afford to spare a few thousand dollars to bail you out.

94%+ of federal cases don't even make it to trial, they end up in plea-bargain agreements, because if you don't agree to trumped up charges, they'll stack charges on you, so that you'll either face 90 years in prison or a year with plea-bargain. a sentence given to murderers and the worst of society, if you lose a trial, or a year if you falsely admit your guilt. losing a non-binding LLM trial could be a requirement for all plea-bargains to avoid this injustice.

Don't even get me started on how utter fecal matter like how you dress, how you comb your hair, your ethnicity, how you sound, your last name, what zip code you find yourself in, the mood of the judge, how hungry the judge is, or their glucose level, how much sleep the judge had. all these factors matter. Juries are even worse, they're a literal coin-toss practically.

I say let LLMs be the first layer of justice, let a human judge turn over their judgement, let justice be swift where possible, without making room for injustice. Allow defendants to choose to wait for a human judge instead if they want. Most I'm sure will take a chance with the LLM, and if that isn't in their favor, nothing changes because they'll now be facing a human judge like they would have otherwise. we can eve talk about sealing the details of the LLM's judgement while appeals are in progress to avoid biasing appellate judges and juries.

Or.. you know.. we could dispense with jail? If cops think someone needs to be placed under arrest, they should prove to a judge within 12 hours that the person is a danger to the community. if they're not a danger, ankle monitors should be placed on them, with no restriction on their movement so long as they remain in the jurisdiction. or house-arrest for serious charges. violating terms would mean actual jail. If you don't like LLMs, I hope you support this instead at the very least. The current system is an abomination and an utter perversion of justice.

I'd prefer caning like they do in Singapore and few other places. brutal, but swift, and you can get back to your life without the cruel bureaucracy destroying or murdering you.

Re: GPT-5 outperforms federal judges in legal reasoning experiment

#113
post #87

IANAL, but this seems like an odd test to me. Judges do what their name implies - make judgment calls. I find it re-assuring that judges get different answers under different scenarios, because it means they are listening and making judgment calls. If LLMs give only one answer, no matter what nuances are at play, that sounds like they are failing to judge and instead are diminishing the thought process down to black-…

> Judges do what their name implies - make judgment calls. I find it re-assuring that judges get different answers under different scenarios, because it means they are listening and making judgment calls. I disagree - law should be the same for everyone. Yes sometimes crimes have mitigating curcumstances and those should be taken into account. However that seems like a separate question of what is and is not illegal.

> law should be the same for everyone

Nah. Too often their "crimes" are actually basic freedoms that they just find it profitable to deny. So many laws are bought and paid for by corporations. There is no need to respect them or even recognize them as legitimate, let alone make them universal.

Re: GPT-5 outperforms federal judges in legal reasoning experiment

#114
post #110

Earlier quoted context omitted.

Yeah, I'm reminded of the various child porn cases where the "perpetrator" is a stupid teenager who took nude pics of themselves and sent them to their boy/girlfriend. Many of those cases have been struck down by judges because the letter of the law creates a non-sequitur where the teenager is somehow a felon child predator who solely preyed on themselves, and sending them to jail and forcing them to sign up for a se…

This example feels more like a bug in the law itself that should be corrected. If this behavior is acceptable then it should be legal so we can avoid everyone the hassle in the first place. I bet AI would be great at finding and fixing these bugs.

AI would be great IF they know what to find

The state of current AI does not give them ability to know that, so the consideration is likely to be dropped

Re: GPT-5 outperforms federal judges in legal reasoning experiment

#116

Earlier quoted context omitted.

A major component of being a judge is to be objective, given the facts.

Yes, but whether they admit it or not, as humans subjectivity, whether informed by culture, opinion, experience, etc, creeps in. There's also variation in how a judge applies objective assessment to law; my interpretation of law may be different to someone else's.

I think we are discussing separate things. I'm talking about requirements and you're talking implementation. The requirement is for judges to be objective and impartial. Turns out AI does a better job of implementing the requirement than humans.

Re: GPT-5 outperforms federal judges in legal reasoning experiment

#117

I wonder whether the original study was in GPT-5's training data. I asked it whether this was the case, and it denied it, but I have no idea whether that result is credible.

I was also wondering this, and in one of the footnotes they say "Given that our experiment was conducted in 2025, one might wonder whether Kansas’ updated law is reflected in GPT’s training data and thus skews its decisions. We find no evidence of such contamination." when talking about a specific updated law. But how does one have 'no evidence of such contamination' without seeing the training data?

They have no evidence of such contamination, not evidence of no such contamination

Re: GPT-5 outperforms federal judges in legal reasoning experiment

#118
post #110

Earlier quoted context omitted.

Yeah, I'm reminded of the various child porn cases where the "perpetrator" is a stupid teenager who took nude pics of themselves and sent them to their boy/girlfriend. Many of those cases have been struck down by judges because the letter of the law creates a non-sequitur where the teenager is somehow a felon child predator who solely preyed on themselves, and sending them to jail and forcing them to sign up for a se…

This example feels more like a bug in the law itself that should be corrected. If this behavior is acceptable then it should be legal so we can avoid everyone the hassle in the first place. I bet AI would be great at finding and fixing these bugs.

now you are talking about replacing not judges, but your elected representatives.

Re: GPT-5 outperforms federal judges in legal reasoning experiment

#119
post #110

Earlier quoted context omitted.

Yeah, I'm reminded of the various child porn cases where the "perpetrator" is a stupid teenager who took nude pics of themselves and sent them to their boy/girlfriend. Many of those cases have been struck down by judges because the letter of the law creates a non-sequitur where the teenager is somehow a felon child predator who solely preyed on themselves, and sending them to jail and forcing them to sign up for a se…

This example feels more like a bug in the law itself that should be corrected. If this behavior is acceptable then it should be legal so we can avoid everyone the hassle in the first place. I bet AI would be great at finding and fixing these bugs.

There are no “bugs” in human institutions like law. There are always going to be edge cases and nuances that require a human to evaluate.
Post reply on HN