Live data from Hacker News

AI outperforms law professors in Stanford Law study

law.stanford.edu

341–350 of 384 posts

Re: AI outperforms law professors in Stanford Law study

#341
post #115

Earlier quoted context omitted.

Well this is largely the fault of law itself. especially english style law. A legal, parseable code, in which not every single tiny municipality (some less than 1 square mile) has their own set of rules and laws, not all published or available - but which citizens are expected to abide by of course - how could we expect AI to do well and not some typical TV southern lawyer who knows the judge?

I could not agree more. A simple example: it boggles my mind how every state organizes their statutes in entirely dissimilar ways. I'm not sure there's a need for every state to have slightly different wording for a murder statute in the first place, but even assuming there is, why do they all have to be scattered around in different code sections instead of every state just following some consistent convention like…

[dead]

Re: AI outperforms law professors in Stanford Law study

#342

Earlier quoted context omitted.

Sure, but in two years AI has gone from “impressive tool, but not a replacement for knowledge workers” to “the study where it beats our highest caliber of knowledge workers may have some methodological deficits.” In another two years it’s going to be curtains.

Autopilots have been able to land planes for years (decades?), and yet they still don't land passengers planes at any increased rate.

[deleted]

Re: AI outperforms law professors in Stanford Law study

#343

I find this study quite suspect. I'd have to dive deeper but there's definitely significant alarm bells that should be going off for anyone reading. Figure 2 (page 6) screams problems. There's only 16 professors (3k comparisons each?!?!) and the professors are all over the place. That's very high variance, suggesting the study has no meaningful statistical power. Poor instructor 16 can't catch a break lol There's als…

> confidently say this paper has a smell to it. I can't call it bullshit, but there are red flags all over

You can confidently say that you are unsure?

Re: AI outperforms law professors in Stanford Law study

#344
post #80

Earlier quoted context omitted.

But can an LLM come up with questions like what the definition of is is? Seems to me there's a lot of "depends on how you read it" type of stuff that lawyers excel at finding novel interpretations. So what coders thinking of as rules are much less straight forward to understand when it comes to laws

I think that’s a different task than the one OP is referring to. To your example, I’m not familiar with the capability of LLMs in that regard. I have struggled with using the AI features of westlaw when it comes to that sort of argument. (Basically, making an argument that strays from typical route, because that’s the position you happen to find yourself representing.)

I'd only be guessing, but I'd imagine that trying to simulate being a lawyer for someone trying to do something shady would really push an LLM. Imagine being a lawyer for Trump. Could it ever come up with the arguments that his lawyers have? God help us all if they do

Re: AI outperforms law professors in Stanford Law study

#345

I find this study quite suspect. I'd have to dive deeper but there's definitely significant alarm bells that should be going off for anyone reading. Figure 2 (page 6) screams problems. There's only 16 professors (3k comparisons each?!?!) and the professors are all over the place. That's very high variance, suggesting the study has no meaningful statistical power. Poor instructor 16 can't catch a break lol There's als…

I never get the same answer from any two lawyers. I hate law as a result. With developers you might get disagreements based on experience, but there's usually a strong consensus on specific things, with lawyers and courts its all over the flipping place. I wouldn't be surprised if LLMs can "pass" on paper (ie college exams) but in practice, they might 'struggle' in different courts. ...On the other hand, if an LLM ha…

I now foresee a future where law firms have models trained on all the transcriptions of individual judges, lawyers and prosecutors, and run agents against them to decide on the optimal strategy for a case.

Re: AI outperforms law professors in Stanford Law study

#346

I find this study quite suspect. I'd have to dive deeper but there's definitely significant alarm bells that should be going off for anyone reading. Figure 2 (page 6) screams problems. There's only 16 professors (3k comparisons each?!?!) and the professors are all over the place. That's very high variance, suggesting the study has no meaningful statistical power. Poor instructor 16 can't catch a break lol There's als…

More than that, the entire structure of the study is pointless. They set up as a question/response and then had humans rate the response. That's literally what LLM's are trained to do, which ultimately is convincing a human to click the "I like this one better" button on it's response.

LLMs are trained to convince a typical human to click the "I like this one better" on their response.

Convincing a human law professor to click the "I would prefer to deliver this response to a student" button, and to not click the "this response is pedagogically harmful" button is a different task!

I could imagine an LLM convincing a typical human to click the "I like this one better" button with flattery, or with nice-sounding platitudes, or with hand-wavey explanations that sound plausible. And in fact that's exactly what LLMs do when they go wrong - they bluff and output superficially plausible nonsense!

But these weren't typical humans, these were law professors specifically tasked with deciding which response was a better option to give to students as a canonical answer to a contract law question. So I think this is a genuinely impressive result.

Re: AI outperforms law professors in Stanford Law study

#347
I wonder if this could be explained in a similar way to Hollywood movies. If the movies are designed to please the largest group of people, there is a greater chance people will choose to see it than another movie. The human law professors come with their own personalities, beliefs, and opinions that come through in their writing. An LLM has been trained to please the largest swathe of the population. That doesn't mean the answer is better; just like Captain America isn't necessarily better than American Beauty.

Re: AI outperforms law professors in Stanford Law study

#348
Figure I.1 is telling. It shows answer length is the strongest predictor of win rate. I suspect this is due to the flawed methodology of the study. Professors were instructed to be succinct ("Please be concise. We expect that each answer takes no more than 3 minutes to write down.") and likely erred on the short side. Also, professors may not have put great effort into their written answers, especially when already trying to be concise. This isn't the headline the authors think it is.

Re: AI outperforms law professors in Stanford Law study

#349
post #94

Earlier quoted context omitted.

A human has a motive that exists that frames the thought being expressed. An LLM is going to be creating a “de novo” thought in response to a line of questioning.

Psychology has shown that a lot of those motives are just post hoc narratives, similar to LLM.

Or, as the extreme claim (and the one that I believe), all of them are: https://en.wikipedia.org/wiki/Epiphenomenalism

Re: AI outperforms law professors in Stanford Law study

#350

I find this study quite suspect. I'd have to dive deeper but there's definitely significant alarm bells that should be going off for anyone reading. Figure 2 (page 6) screams problems. There's only 16 professors (3k comparisons each?!?!) and the professors are all over the place. That's very high variance, suggesting the study has no meaningful statistical power. Poor instructor 16 can't catch a break lol There's als…

Sure, but in two years AI has gone from “impressive tool, but not a replacement for knowledge workers” to “the study where it beats our highest caliber of knowledge workers may have some methodological deficits.” In another two years it’s going to be curtains.

I'd say if it does have methodological deficits, it should be ignored. Measuring a length with a wet spaghetti can only result in nonsense.
Post reply on HN