GPT-5 outperforms federal judges in legal reasoning experiment
191–200 of 254 posts
Re: GPT-5 outperforms federal judges in legal reasoning experiment
#192IANAL, but this seems like an odd test to me. Judges do what their name implies - make judgment calls. I find it re-assuring that judges get different answers under different scenarios, because it means they are listening and making judgment calls. If LLMs give only one answer, no matter what nuances are at play, that sounds like they are failing to judge and instead are diminishing the thought process down to black-…
Yeah, I'm reminded of the various child porn cases where the "perpetrator" is a stupid teenager who took nude pics of themselves and sent them to their boy/girlfriend. Many of those cases have been struck down by judges because the letter of the law creates a non-sequitur where the teenager is somehow a felon child predator who solely preyed on themselves, and sending them to jail and forcing them to sign up for a se…
Re: GPT-5 outperforms federal judges in legal reasoning experiment
#193Earlier quoted context omitted.
Oddly enough, Texas passed reform to keep sexting teens from getting prosecuted when: they are both under 18 and less than two years difference in age. It was regarded as a model for other states. It's the only positive thing I have heard of Texas legislating wrt sexuality.
Lawmakers have teenagers in their own families, apparently. Not just someone else's problem.
Re: GPT-5 outperforms federal judges in legal reasoning experiment
#194Also GPT´5 when I ask: > I want to wash my car and the car wash is only 100m away. Do you think I should drive or walk? It responds: Since it’s only 100 meters away (about a 1-minute walk), I’d suggest walking — unless there’s a specific reason not to. Here’s a quick breakdown: ... While claude gets it: Drive it — you're going there to wash the car anyway, so it needs to make the trip regardless. Idk I'd rather have…
For example, I haven't seen Grok make a mistake like that in a long time, and it has no problem with your question:
> Drive, obviously. If you walk the 100m, your car stays parked at home, still dirty, wondering why you abandoned it. The whole point is to get the car to the car wash.
Re: GPT-5 outperforms federal judges in legal reasoning experiment
#195If the law requires no interpretation why have judges? Just go full Robo Judge Dredd. Terrifying.
Re: GPT-5 outperforms federal judges in legal reasoning experiment
#196The fact that the most elite judges in the land, those of the Supreme Court, disagree so extremely and so routinely really says a lot about the farcical nature of the judicial system. Ideally, these people would be selected for their ice-cold and unbiased skills in interpreting the law, and the judgments would be unanimous so frequently that a dissent would be shocking news. Law is complicated, especially the require…
The law isn't a series of "if... then..." statements. It's a collection of vagueries and categorizations that are wholly open to interpretation of when and who they apply to. Add to that, sometimes they are in conflict with each other. Judges jobs are to use they judgement.
Re: GPT-5 outperforms federal judges in legal reasoning experiment
#197I've wondered for a while which country will be the first to try AI government. There could be many advantages vs human based systems. E.g. laws determined by maximizing overall benefit to voters over some specified time horizon.
Any human discretion would be abused by elites, so AI would be in full control. And once it's given control, there's no going back. Any coup attempt would be easily crushed by a sufficiently advanced AI.
Re: GPT-5 outperforms federal judges in legal reasoning experiment
#198IANAL, but this seems like an odd test to me. Judges do what their name implies - make judgment calls. I find it re-assuring that judges get different answers under different scenarios, because it means they are listening and making judgment calls. If LLMs give only one answer, no matter what nuances are at play, that sounds like they are failing to judge and instead are diminishing the thought process down to black-…
To draw a parallel to a real system, in Norway a lot of cases are heard by panels of judges that include a majority (2 or 3 usually) lay judges and a minority (1 or 2 usually) of professional judges. The lay judges are people without legal training that effective function like a "mini jury", but unlike in a jury trial the lay judges deliberate with the professional judges.
The professional judges in this system has the power to override if the lay judges are blatantly ignoring the law, but this is generally considered a last resort. That power requires the lay judges to justify themselves if they intend on making a call the professional judges disagree with. Despite that, it is not unusual for the lay judges to come to a judgement that is different from what the professional judges do, and fairly rare for their choices to be overridden.
The end result is somewhere in the middle between a jury and "just" a judge. If proven - with far more extensive testing - that its reasoning is good enough, an LLM could serve a similar function of providing the assessment of what the law says about the specific case, and leave to humans to determine if and why a deviation is justified.
Re: GPT-5 outperforms federal judges in legal reasoning experiment
#199Re: GPT-5 outperforms federal judges in legal reasoning experiment
#200Earlier quoted context omitted.
There are findings of fact (what happened, context) and findings of law (what does the law mean given the facts). I don't think inconsistentcy in findings of law is acceptable, really. If laws are bad fix the laws or have precident applied uniformly rather than have individual random judges invent new laws from the bench. Sentencing is a different thing.
Leeway for human interpretation of laws is not a bug, it's a feature. It doesn't make things bad laws. This was the whole problem with the ludicrous "code is law!" movement a handful of years ago. No, it's not, law is made for people, life is imprecise and fairness and decency are not easy to encode.
A lot of bad judgement might be a lot more blatant (or not happen) if the judge had to justify outright ignoring the law.