Live data from Hacker News

GPT-5 outperforms federal judges in legal reasoning experiment

papers.ssrn.com

231–240 of 254 posts

Re: GPT-5 outperforms federal judges in legal reasoning experiment

#231

Earlier quoted context omitted.

The law isn't a series of "if... then..." statements. It's a collection of vagueries and categorizations that are wholly open to interpretation of when and who they apply to. Add to that, sometimes they are in conflict with each other. Judges jobs are to use they judgement.

It's not currently, but if we were able to use AI to generate laws in an objective and logically sound way based on general principles like "don't harm others or their property", we'd be much better off.

> if we were able to use AI to generate laws in an objective and logically sound way ... we'd be much better off.

A major role of judges is specifically to not do that because there are circumstances that will not have been thought of at the time of a law being written, new laws will be written that interact in unforeseen ways with existing laws and/or common views on justice can change over time.

It may be technically illegal to destroy a person's property but no judge is going to convict someone who breaks down a person's front door because they heard someone crying for help inside. That's a simple example but there would have to be enumerable exceptions to every single law for an objective/logical AI to do justice.

Rather than try to enumerate the enumerable, we let judges judge.

Re: GPT-5 outperforms federal judges in legal reasoning experiment

#232

Earlier quoted context omitted.

The law isn't a series of "if... then..." statements. It's a collection of vagueries and categorizations that are wholly open to interpretation of when and who they apply to. Add to that, sometimes they are in conflict with each other. Judges jobs are to use they judgement.

> The law isn't a series of "if... then..." statements I mean, it's literally called (in the US, at least) the United States Code [1]. [1] https://en.wikipedia.org/wiki/United_States_Code

And everyone knows that, if something is codified, it never has any unintended consequences.

Re: GPT-5 outperforms federal judges in legal reasoning experiment

#233
Right... So how far are we from creating SIBYL? Watch that show (Psycho-Pass) to see what happens when AI starts judging humanity at scale and who is behind that "AI".

More to the point, this decade is going to set some scary precedents that would need to be overturned. Would AI know which case law carries more weight and which was purely politically motivated with no basis in reality?

Re: GPT-5 outperforms federal judges in legal reasoning experiment

#234
post #91

Tim & Eric: In our 2009 sketch we invented Cinco e-Trial as a cautionary tale. Tech Company: At long last, we have created Cinco e-Trial from classic sketch "Don't Create Cinco e-Trial" https://www.youtube.com/watch?v=vKety3N00Gk

The Great Job! Cinco skits weren't usually cautionary tales, but parodies of how product marketing overlaps with mundane reality. E-Trial, My New Pep-Pep and Cinco-fone are all devoid of any moral lesson. They're real infomercials for fake products, which hammers home how harmful and deluded unregulated advertisement has gotten in 2026.

Re: GPT-5 outperforms federal judges in legal reasoning experiment

#235

Also GPT´5 when I ask: > I want to wash my car and the car wash is only 100m away. Do you think I should drive or walk? It responds: Since it’s only 100 meters away (about a 1-minute walk), I’d suggest walking — unless there’s a specific reason not to. Here’s a quick breakdown: ... While claude gets it: Drive it — you're going there to wash the car anyway, so it needs to make the trip regardless. Idk I'd rather have…

Strange because I used your exact prompt and GPT 5 gave the correct answer and immediately explained how the question was maliciously constructed.

GPT4o was duped though.

Re: GPT-5 outperforms federal judges in legal reasoning experiment

#236
post #47

Frankly I don’t care, I’ll take human judges any day, because they have something AI does not: flesh and bone and real skin in the game.

Real skin in the game is also known as bias. That's an example of something a judge should not have.

Judges should have some amount of biases. A cold, calculating, unbiased judge in a world where laws are written by fallible humans would be terrible.

Judges should be able to apply judgement, not be merely automatons that sentence according to only the exact letter of the law.

Laws are not perfect, we need human judges.

Finally, if we are to submit ourselves to judgement by others, I gives me some comfort to know that the being judging me is equally mortal and can be deposed if necessary, as they are flesh and blood like me.

Re: GPT-5 outperforms federal judges in legal reasoning experiment

#237
post #108
post #47

Frankly I don’t care, I’ll take human judges any day, because they have something AI does not: flesh and bone and real skin in the game.

Not really. Ultimately it's just a job and a job without any tangible benefit to doing well. Most regular folk that end up in front of a judge would do well to have a quick and predictable decision. It's months to years before things happen in court and are usually gated behind 10s of thousands in legal fees or a ton of effort. To have a judge bot available for a decision immediately is enormously beneficial.

Sounds to me like we’re bringing too many people before judges then.

Re: GPT-5 outperforms federal judges in legal reasoning experiment

#238

Earlier quoted context omitted.

Laws are written to be interpreted and applied by humans. They aren’t computer programs. They are full of ambiguity. Much of this is by design because there are too many possible edge cases to design a fully algorithmic unambiguous legal system.

True, but its not a free for all. Judges (especially in a common law juridsiction) are supposed to be consistent and interpret laws following certain principles. There are more right and less right interpretations - thus we can grade judges on how well they do their job.

Nothing at all in the post you replied to

> different answers under different scenarios

implies that anyone is inconsistently applying any principles.

Re: GPT-5 outperforms federal judges in legal reasoning experiment

#239
post #146

Earlier quoted context omitted.

You might have misread it. Texas' model is decriminalizing teens sexting, not criminalizing it

I think they refer to the fact that, exposed as GP did, looks like there is a loophole if 2 teenagers started their relationship at 17 and 15, and once they become 18 and 16, sexting is suddenly illegal.

Best guess without looking it up, they meant to say "either" and not "both"; these are called "Romeo and Juliet laws" and are nothing new.

Re: GPT-5 outperforms federal judges in legal reasoning experiment

#240

Earlier quoted context omitted.

> to appear just to people. The best way to appear just is to be just. But I'm not sure what your argument is. It is our duty as citizens to encourage the system to be just. Since there is no concrete mathematical objective definition of justice, well, then... all we can work with is the appearance. So I don't think your insight is so much based on some diabolical deep state thinking but more on the limitations of pr…

It also involves being able to look the person doing the sentencing in the eye and them telling you their reasons for the ruling to the face; being able to argue and present evidence in front of a neutral arbiter. Facebook's moderation might well be just but it lacks the accountability and openness and humanity to appear just. (It also isn't just but I'm saying that even if it was, it would not appear so.)

I agree. I say more here[0] and think the inability to face your accusers is not only unconstitutional but inhumane.

[0] https://news.ycombinator.com/item?id=46984602

Post reply on HN