Live data from Hacker News

GPT-5 outperforms federal judges in legal reasoning experiment

papers.ssrn.com

171–180 of 254 posts

Re: GPT-5 outperforms federal judges in legal reasoning experiment

#171
post #110

Earlier quoted context omitted.

Yeah, I'm reminded of the various child porn cases where the "perpetrator" is a stupid teenager who took nude pics of themselves and sent them to their boy/girlfriend. Many of those cases have been struck down by judges because the letter of the law creates a non-sequitur where the teenager is somehow a felon child predator who solely preyed on themselves, and sending them to jail and forcing them to sign up for a se…

This example feels more like a bug in the law itself that should be corrected. If this behavior is acceptable then it should be legal so we can avoid everyone the hassle in the first place. I bet AI would be great at finding and fixing these bugs.

> If this behavior is acceptable then it should be legal so we can avoid everyone the hassle in the first place.

Codifying what is morally acceptable into definitive rules has been something humanity has struggled with for likely much longer than written memory. Also while you're out there "fixing bugs" - millions of them and one-by-one - people are affected by them.

> I bet AI would be great at finding and fixing these bugs.

Ae we really going to outsource morality to an unfeeling machine that is trained to behave like an exclusive club of people want it to?

If that was one's goal, that's one way to stealthily nudge and undermine a democracy I suppose.

Re: GPT-5 outperforms federal judges in legal reasoning experiment

#173

IANAL, but this seems like an odd test to me. Judges do what their name implies - make judgment calls. I find it re-assuring that judges get different answers under different scenarios, because it means they are listening and making judgment calls. If LLMs give only one answer, no matter what nuances are at play, that sounds like they are failing to judge and instead are diminishing the thought process down to black-…

Disagree completely. Judgement of the sort you're describing should be done at the legislative phase (i.e. writing code).

Inconsistent execution/application of the law is how bias happens. If a judgement done to the letter of the law feels unjust to you, change the letter of the law.

Re: GPT-5 outperforms federal judges in legal reasoning experiment

#174
post #146

Earlier quoted context omitted.

> It was regarded as a model for other states. Really? That "model" has the common, but obviously extremely undesirable, feature of criminalizing sexual relationships between students in the same grade that were legal when they formed . How could it be regarded as a model for anyone else?

You might have misread it. Texas' model is decriminalizing teens sexting, not criminalizing it

I didn't misread it, but apparently you did.

Why is criminalizing an existing legal relationship a good idea?

Re: GPT-5 outperforms federal judges in legal reasoning experiment

#175
post #87

Earlier quoted context omitted.

> Judges do what their name implies - make judgment calls. I find it re-assuring that judges get different answers under different scenarios, because it means they are listening and making judgment calls. I disagree - law should be the same for everyone. Yes sometimes crimes have mitigating curcumstances and those should be taken into account. However that seems like a separate question of what is and is not illegal.

Laws are written to be interpreted and applied by humans. They aren’t computer programs. They are full of ambiguity. Much of this is by design because there are too many possible edge cases to design a fully algorithmic unambiguous legal system.

True, but its not a free for all. Judges (especially in a common law juridsiction) are supposed to be consistent and interpret laws following certain principles. There are more right and less right interpretations - thus we can grade judges on how well they do their job.

Re: GPT-5 outperforms federal judges in legal reasoning experiment

#176

Earlier quoted context omitted.

The reference to vigilante justice may be about killing a suspect before they're imprisoned or even tried, such as when a mob storms the local jail. The theory is, if people believe only death can bring justice, and the state doesn't have the death penalty, then the vigilantes will take matters into their own hands. Ergo, the state should have the death penalty. Having recently done an in-depth review of arguments fo…

I see; this makes more sense. It's a little hard to imagine these days though, but ages ago, mobs storming the local jail and hanging a suspect wasn't that uncommon.

> ages ago, mobs storming the local jail and hanging a suspect wasn't that uncommon.

Sometimes, suspects don't even make it to the jail.

https://en.wikipedia.org/wiki/Jack_Ruby_Shoots_Lee_Harvey_Os...

https://en.wikipedia.org/wiki/Tulsa_race_massacre

Uncommon or not, vigilantism is incompatible with justice on a societal level, regardless of any alleged guilt of offenders.

Without a showing of evidence, a trial of the accused, and a verdict that withstands judgment, we're left with theories and conjecture, and hatchets long left unburied.

Re: GPT-5 outperforms federal judges in legal reasoning experiment

#177

IANAL, but this seems like an odd test to me. Judges do what their name implies - make judgment calls. I find it re-assuring that judges get different answers under different scenarios, because it means they are listening and making judgment calls. If LLMs give only one answer, no matter what nuances are at play, that sounds like they are failing to judge and instead are diminishing the thought process down to black-…

Yeah, I'm reminded of the various child porn cases where the "perpetrator" is a stupid teenager who took nude pics of themselves and sent them to their boy/girlfriend. Many of those cases have been struck down by judges because the letter of the law creates a non-sequitur where the teenager is somehow a felon child predator who solely preyed on themselves, and sending them to jail and forcing them to sign up for a se…

Man, this is one of the ways society has fundamentally broken - all the 'think of the children' arguments, resting on the belief that children are so sacred, that any sort of leinency or consideration of circumstances is forbidden - lest someone guilty of molesting them might walk free.

Well now we know for a fact that some of the people making these arguments very thinking of the children very much.

Re: GPT-5 outperforms federal judges in legal reasoning experiment

#178

Earlier quoted context omitted.

Sorry but that seems like an insane system where whole classes of actions effectively are illegal but probably okay if you're likeable. In your scenario the obvious solution is to amend the law and pardon people convinced under it. B/c what really happens is that if you have a pretty face and big tits you get out of speeding tickets b/c "gosh well the law wasn't intended for nice people like you"

It isn't "my scenario". These are real cases. https://www.aclu-mn.org/press-releases/victory-judge-dismiss... "In his decision, Judge Cajacob asserts that the purpose and intent of Minnesota’s child pornography statute does not support punishing Jane Doe for explicit images of herself and doing so “produces an absurd, unreasonable, and unjust result that utterly confounds the statue’s stated purpose.”" Nothing in the…

> It isn't "my scenario". These are real cases

maybe English isnt your native language, but "scenario" doesnt require the situation to be not real

> Nothing in there about "likeability" or "we let her off because she had nice tits"

We have no way to know if likeability played in to it. When rules are bendable then they are bent to the likeable and attractive. My example of a traffic stop is analogous and more directly relatable

> This is how the law has always worked, and if you've thought otherwise then consider you've been living under this "insane system" for your entire life

You seem to have some reading comprehension issues.. I never suggested its not currently working that way and i never suggested the current situation is not insane. If you think the current system is sane and great then thats your opinion

Everyone i know whos had to deal with the US legal system has only related horror stories

Re: GPT-5 outperforms federal judges in legal reasoning experiment

#179

Earlier quoted context omitted.

The parent comment present a scenario where the law is ignored b/c the judge decides for himself it shouldn't apply. I'm pointing out that this kind of approach is fundamentally unjust and wrong. "And sure you can say the laws should be written better, but so long as the laws are written by humans that will simply not be the case" The obvious solution is dismissed

Are you a bot? Your name is contrarian1234 and you lack sophisticated interpretations of statements.

given your inability to engage with an opposing point of view, youre definitely not a bot. So ill take your ad hominem as praise

Re: GPT-5 outperforms federal judges in legal reasoning experiment

#180

I was diagnosed with a rare blood disease called Essential Thrombocythemia (ET) which is part of a group of diseases called myeloproliferative neoplasms. This happened about three years ago. Recently, I decided to get a second opinion and my new specialist changed my diagnosis from ET to Polycythemia Vera (PV). She also highly recommended I quickly go and give blood to lower my haematocrit levels as it put me at a mu…

I think AI "slop" will improve medical diagnoses dramatically. Let's assume for a second that the first specialist did not graduate at the top of their class.

The year is 2030, when LLMs are more pervasive. The first specialist now asks you to wait, heads into the other room and double-checks their ET diagnosis with AI. Doing so has become standard practice to avoid malpractice suits. The model persuades them to diagnose PV, avoiding a Type-II error.

But let's say the model gets it wrong too. You eventually visit the second specialist, who did graduate at the top of their class. The model says ET, but the specialist is smart enough to tell that the model is wrong. There is some risk that the second specialist takes the CYA route, but I'd expect them not to. They diagnose PV, avoiding a Type-I error.

Post reply on HN