Live data from Hacker News

AI outperforms law professors in Stanford Law study

law.stanford.edu

101–110 of 384 posts

Re: AI outperforms law professors in Stanford Law study

#101
I find this study quite suspect. I'd have to dive deeper but there's definitely significant alarm bells that should be going off for anyone reading.

Figure 2 (page 6) screams problems. There's only 16 professors (3k comparisons each?!?!) and the professors are all over the place. That's very high variance, suggesting the study has no meaningful statistical power. Poor instructor 16 can't catch a break lol

There's also really clear bias given that the main results only feature Google models. Other models show up elsewhere, why not there?

I'm no lawyer, but I'm a pretty competent statistician and can confidently say this paper has a smell to it. I can't call it bullshit, but there are red flags all over

Re: AI outperforms law professors in Stanford Law study

#102
post #7

As a software engineer I have some intuition for what the risks are of letting agents do some tasks vs others. I don't have a similar intuition calibrated for what could go wrong when asking AI to draft a legal document. Some things seem harmless, i.e. drafting a will, but I don't really know- our legal system is notoriously rife with footguns.

I'm afraid since claude cheats in benches, what will it do with law?

Hmm, what’s the law equivalent of using docker to bypass sudo?

Re: AI outperforms law professors in Stanford Law study

#103

What the LLM cannot do is explain why it said what it said, when cross-examined. It simply hallucinates the best account of why someone would have said such a thing as it said, same as it can give a probable account of why someone else said something different. The question 'But why did you say this not that ...?' does not lead it to make explicit its grounds for what it said, but just to make a new more complicated…

Same is probably true of humans. In a conversation, we often respond from instinct, then work backwards to a rationalization only when asked. For more considered thoughts, if we’re lucky, we can remember our “reasoning traces” but that’s as deep as our introspection goes. Unless we’re neuroscientists, we don’t even know how many neurons we have, let alone have any understanding of how they generate our thoughts. Moti…

No, it happens in the immediate context, where e.g. we say 'No I meant Meredith Jones, not Meredith Smith'- and the possibility of this elaboration is actually part of ordinary communication. I did mean Meredith Jones, not Meredith Smith - thus the use of the past tense The LLM will just give the best answer for what one might have meant, completely reopening calculation.

The point is familiar but there are good illustrations in the Atlantic article by a book editor. At first it seems abstract AI hate, but then she gets to the details. AI text cannot be edited. https://www.theatlantic.com/technology/2026/05/how-to-tell-a... or https://archive.ph/YJsGK

Re: AI outperforms law professors in Stanford Law study

#104

Earlier quoted context omitted.

I think this is probably true for most skilled professions. AI is best used in the hands of folks already knowledgeable in the skills/professions they are using it for. I liken it to me googling things as a sysadmin vs. Jane from accounting doing it. The non-tech end user is far more likely to make the problem worse, or install something sketchy from the ad riddled results than I am, or one of my help desk employees…

> I think this is probably true for most skilled professions. I agree, BUT I also find that it's easy for experts to atrophy quickly. When the AI is right 80/90% of the time it lulls you into over confidence. I find those that are best and make the greatest use are the ones who remain skeptical but also use the tool. The same people who were already nuanced and picky before AI. The same people who already doubted and…

I would agree with this point and as I explained in a comment replying to the GP comment above, that atrophy is far more dangerous in the legal field than it is with code because legal documents do not benefit from the structural safeguards available for code, like automated testing, static typing, static analysis tools, etc. IME with legal LLMs so far, they are easily in that most dangerous valley where they can lull you into a false sense of security while still introducing extremely dangerous mistakes that are frequently difficult to detect without very careful reading.

The danger of those mistakes creeping in also grows exponentially the farther a lawyer strays from their core legal expertise. There are a few statutes I know inside and out, and I can spot LLM analytical errors related to them in a split second, but once I venture out into domains where I am not an expert (but where I am nevertheless reasonably qualified to practice), it becomes much harder to spot drafting mistakes because I have not refreshed my own understanding of the law by reviewing the relevant cases or statutes as I would when drafting the analysis myself from scratch.

Re: AI outperforms law professors in Stanford Law study

#105
I'm going to need some legal help for my startup. But I can't pay much. So I figured I will ask AI all relevant questions, as well as forms filled etc. Perhaps even create a patent-application for me.

THEN I find a human lawyer and give AI's answers to them and say "Can you find any errors in this? Can you improve it?" .

That way I think my legal bills should be smaller because the AI has already done most of the work. What do you think? Which LLM is best for legal work?

Re: AI outperforms law professors in Stanford Law study

#106
post #97

What the LLM cannot do is explain why it said what it said, when cross-examined. It simply hallucinates the best account of why someone would have said such a thing as it said, same as it can give a probable account of why someone else said something different. The question 'But why did you say this not that ...?' does not lead it to make explicit its grounds for what it said, but just to make a new more complicated…

LLMs hallucinate, because humans hallucinate. Asking the LLM in a way where it annotates its sources, it can greatly increase the pattern matching to closely simulate logic, just like in humans. I understand the question of why did you say this, not that, I have seen other ways of asking that which do not seem to trigger the LLMs over-response in the other direction.

Humans hallucinate because they take shrooms or have schizophrenia.

Re: AI outperforms law professors in Stanford Law study

#107
post #89

Earlier quoted context omitted.

for what it's worth I have no idea why it would be nonsense to question institutional motivations especially in the context of an academic article that could easily be corporate propaganda, I also think that shutting conversations down is much more harmful than discussing topics that are potentially harmful

Completely unevidenced conspiracy theories can only harm the discussion. The only possible benefit is to disconfirm conspiracy theories and discourage paranoid thinking. The odds that Standford as an institution are astroturfing on HN round down to 0. What they're almost certainly observing is that these critical comments are being flagged as inappropriate. People make inappropriate comments that happen to contain cr…

may I ask why you effectively said 'conversation over due to harm reasons' instead of asking for evidence to support the conclusion that you believe is not possible? I don't see why it is inherently harmful to discuss the seemingly impossible. I also don't see why it's relevant to bring up your n=1 sample (although it is as relevant as my n=1 sample, which has plenty of astroturfing witnessing [unspecific to Stanford])

Re: AI outperforms law professors in Stanford Law study

#108
post #97

What the LLM cannot do is explain why it said what it said, when cross-examined. It simply hallucinates the best account of why someone would have said such a thing as it said, same as it can give a probable account of why someone else said something different. The question 'But why did you say this not that ...?' does not lead it to make explicit its grounds for what it said, but just to make a new more complicated…

LLMs hallucinate, because humans hallucinate. Asking the LLM in a way where it annotates its sources, it can greatly increase the pattern matching to closely simulate logic, just like in humans. I understand the question of why did you say this, not that, I have seen other ways of asking that which do not seem to trigger the LLMs over-response in the other direction.

No, the hallucination of its reasons follows immediately from the technique of probabilistic inference. You can see this in real time, just ask 'why did you use this word, not that word?' It is in the position of a desperate liar. All its responses are essentially 'rationalizations'

Re: AI outperforms law professors in Stanford Law study

#109

I'm going to need some legal help for my startup. But I can't pay much. So I figured I will ask AI all relevant questions, as well as forms filled etc. Perhaps even create a patent-application for me. THEN I find a human lawyer and give AI's answers to them and say "Can you find any errors in this? Can you improve it?" . That way I think my legal bills should be smaller because the AI has already done most of the wor…

i use codex to do initial research and draft texts (in typst). i use files-output skill so that all research contexts are rendered into files md files.

i do second phase on codex, by asking to download all pdfs and extract all text of laws it references. can repeat fully local research step.

after i ask gemini to find issues and criticize.

UPDATE: there many legal skills on github to try, not used so any yet

Re: AI outperforms law professors in Stanford Law study

#110
post #22

Earlier quoted context omitted.

I would think that LLMs would be better at avoiding foot-guns. That’s a situation where you have a list of well known rules and potential pit falls, and the work of the lawyer is to apply those to a fact pattern. That’s something that has been hard to automate programmatically, because the fact patterns are similar but different. LLMs, however, seem to excel at applying general principles to differing fact patterns.

I would categorize this in the "expertise that people internalize but never figure out how to verbalize" department, and that is a department we have no way to teach an LLM because if nobody is writing out those unspoken, subconscious rules then the LLM has nothing to read about them in its training data.

Good point. Same probably applies to code as well, coders much tell us why they write the cde the way they did. And if they have comments in their code, those are highly untrustworthy because noboy fixes comments if the code works.
Post reply on HN