GPT-5 outperforms federal judges in legal reasoning experiment
1–10 of 254 posts
Re: GPT-5 outperforms federal judges in legal reasoning experiment
#2Digging a bit deeper, the actual paper seems to agree: "For the sake of consistency, we define an “error” in the same way that Klerman and Spamann do in their original paper: a departure from the law. Such departures, however, may not always reflect true lawlessness. In particular, when the applicable doctrine is a standard, judges may be exercising the discretion the standard affords to reach a decision different from what a surface-level reading of the doctrine would suggest"
Re: GPT-5 outperforms federal judges in legal reasoning experiment
#3Re: GPT-5 outperforms federal judges in legal reasoning experiment
#4IANAL, but this seems like an odd test to me. Judges do what their name implies - make judgment calls. I find it re-assuring that judges get different answers under different scenarios, because it means they are listening and making judgment calls. If LLMs give only one answer, no matter what nuances are at play, that sounds like they are failing to judge and instead are diminishing the thought process down to black-…
Re: GPT-5 outperforms federal judges in legal reasoning experiment
#5IANAL, but this seems like an odd test to me. Judges do what their name implies - make judgment calls. I find it re-assuring that judges get different answers under different scenarios, because it means they are listening and making judgment calls. If LLMs give only one answer, no matter what nuances are at play, that sounds like they are failing to judge and instead are diminishing the thought process down to black-…
These were technical rulings on matters of jurisdiction, not subjective judgments on fairness.
"The consistency in legal compliance from GPT, irrespective of the selected forum, differs significantly from judges, who were more likely to follow the law under the rule than the standard (though not at a statistically significant level). The judges’ behavior in this experiment is consistent with the conventional wisdom that judges are generally more restrained by rules than they are by standards. Even when judges benefit from rules, however, they make errors while GPT does not.
Re: GPT-5 outperforms federal judges in legal reasoning experiment
#6You can also avoid "hungry judge effect" by making sure GPT is always fully charged before prompting it.
Re: GPT-5 outperforms federal judges in legal reasoning experiment
#7IANAL, but this seems like an odd test to me. Judges do what their name implies - make judgment calls. I find it re-assuring that judges get different answers under different scenarios, because it means they are listening and making judgment calls. If LLMs give only one answer, no matter what nuances are at play, that sounds like they are failing to judge and instead are diminishing the thought process down to black-…
Re: GPT-5 outperforms federal judges in legal reasoning experiment
#8IANAL, but this seems like an odd test to me. Judges do what their name implies - make judgment calls. I find it re-assuring that judges get different answers under different scenarios, because it means they are listening and making judgment calls. If LLMs give only one answer, no matter what nuances are at play, that sounds like they are failing to judge and instead are diminishing the thought process down to black-…
Re: GPT-5 outperforms federal judges in legal reasoning experiment
#9IANAL, but this seems like an odd test to me. Judges do what their name implies - make judgment calls. I find it re-assuring that judges get different answers under different scenarios, because it means they are listening and making judgment calls. If LLMs give only one answer, no matter what nuances are at play, that sounds like they are failing to judge and instead are diminishing the thought process down to black-…
Re: GPT-5 outperforms federal judges in legal reasoning experiment
#10You can also avoid "hungry judge effect" by making sure GPT is always fully charged before prompting it.
Cases aren't ordered randomly. Obvious cases are scheduled at the end of session before breaks.