GPT-5 outperforms federal judges in legal reasoning experiment
11–20 of 254 posts
Re: GPT-5 outperforms federal judges in legal reasoning experiment
#12You can also avoid "hungry judge effect" by making sure GPT is always fully charged before prompting it.
"hungry judge effect" is a debunked myth.
Re: GPT-5 outperforms federal judges in legal reasoning experiment
#13Re: GPT-5 outperforms federal judges in legal reasoning experiment
#14From the paper:
“we find that the LLM adheres to the legally correct outcome significantly more often than human judges”
That presupposes that a “legally correct” outcome exists
The Common Law, which is the foundation of federal law and the law of 49/50 states, is a “bottom up” legal system.
Legal principals flow from the specific to the general. That is, judges decided specific cases based on the merits of that individual case. General principles are derived from lots of specific examples.
This is different from the Civil Law used in most of Europe, which is top-down. Rulings in specific cases are derived from statutory principles.
In the US system, there isn’t really a “correct legal outcome”.
Common Law heavily relies on “Juris Prudence”. That is, we have a system that defers to the opinions of “important people”.
So, there isn’t a “correct” legal outcome.
Re: GPT-5 outperforms federal judges in legal reasoning experiment
#15Re: GPT-5 outperforms federal judges in legal reasoning experiment
#16Re: GPT-5 outperforms federal judges in legal reasoning experiment
#17The ability of ai to serve as impartial mediators could become the greatest civil rights advance in modern history.
Re: GPT-5 outperforms federal judges in legal reasoning experiment
#18The premise seems flawed. From the paper: “we find that the LLM adheres to the legally correct outcome significantly more often than human judges” That presupposes that a “legally correct” outcome exists The Common Law, which is the foundation of federal law and the law of 49/50 states, is a “bottom up” legal system. Legal principals flow from the specific to the general. That is, judges decided specific cases based…
Remember the article that described LLMs as lossy compression and warned that if LLM output dominated the training set, it would lead to accumulated lossiness? Like a jpeg of a jpeg
Re: GPT-5 outperforms federal judges in legal reasoning experiment
#19IANAL, but this seems like an odd test to me. Judges do what their name implies - make judgment calls. I find it re-assuring that judges get different answers under different scenarios, because it means they are listening and making judgment calls. If LLMs give only one answer, no matter what nuances are at play, that sounds like they are failing to judge and instead are diminishing the thought process down to black-…
I don't trust AI in its current form to make that sort of distinction. And sure you can say the laws should be written better, but so long as the laws are written by humans that will simply not be the case.
Re: GPT-5 outperforms federal judges in legal reasoning experiment
#20The premise seems flawed. From the paper: “we find that the LLM adheres to the legally correct outcome significantly more often than human judges” That presupposes that a “legally correct” outcome exists The Common Law, which is the foundation of federal law and the law of 49/50 states, is a “bottom up” legal system. Legal principals flow from the specific to the general. That is, judges decided specific cases based…