Live data from Hacker News

Re-Evaluating GPT-4's Bar Exam Performance

link.springer.com

1–10 of 139 posts

Re: Re-Evaluating GPT-4's Bar Exam Performance

#2
Very interesting. The abstract claims that although GPT-4 was claimed to score in the 92nd percentile on the bar exam, when correcting for a bunch of things they find that these results are overinflated, and that it only scores in the 15th percentile specifically on essays when compared to only people that passed the bar.

That still does put it into bar-passing territory, though, since it still scores better than about one sixth of the people that passed the exam.

Re: Re-Evaluating GPT-4's Bar Exam Performance

#4
post #2

Very interesting. The abstract claims that although GPT-4 was claimed to score in the 92nd percentile on the bar exam, when correcting for a bunch of things they find that these results are overinflated, and that it only scores in the 15th percentile specifically on essays when compared to only people that passed the bar. That still does put it into bar-passing territory, though, since it still scores better than abo…

If I understand currently, they measured it at the 69th percentile for the full test across all test takers, so definitely still impressive.

Re: Re-Evaluating GPT-4's Bar Exam Performance

#5
A basic problem with evaluations like these is that the test is designed to discriminate between humans who would make good lawyers and humans who would not make good lawyers. The test is not necessarily any good at telling whether a non-human would make a good lawyer, since it will not test anything that pretty much all humans know, but non-humans may not.

For example, I doubt that it asks whether, for a person of average wealth and income, a $1000 fine is a more or less severe punishment than a month in jail.

Re: Re-Evaluating GPT-4's Bar Exam Performance

#6
The bigger issue here is that actual legal practice looks nothing like the bar, so whether or not an llm passes says nothing about how llms will impact the legal field.

Passing the bar should not be understood to mean "can successfully perform legal tasks."

Re: Re-Evaluating GPT-4's Bar Exam Performance

#7
They originally scored against a test usually taken by people who failed the bar.

So, GPT-4 scores closer to the bottom of people who pass the bar the first time. In other words, it matches the people who cull the rules from texts already written, but who cannot apply it imaginatively.

Re: Re-Evaluating GPT-4's Bar Exam Performance

#8

A basic problem with evaluations like these is that the test is designed to discriminate between humans who would make good lawyers and humans who would not make good lawyers. The test is not necessarily any good at telling whether a non-human would make a good lawyer, since it will not test anything that pretty much all humans know, but non-humans may not. For example, I doubt that it asks whether, for a person of a…

Honestly, this is giving the bar exam (and GPT-4) too much credit. The bar tests memorization because it's challenging for humans and easy to score objectively. But memorization isn't that important in legal practice; analysis is. LLMs are superhuman at memorization but terrible at analysis.

Re: Re-Evaluating GPT-4's Bar Exam Performance

#9

A basic problem with evaluations like these is that the test is designed to discriminate between humans who would make good lawyers and humans who would not make good lawyers. The test is not necessarily any good at telling whether a non-human would make a good lawyer, since it will not test anything that pretty much all humans know, but non-humans may not. For example, I doubt that it asks whether, for a person of a…

Honestly, this is giving the bar exam (and GPT-4) too much credit. The bar tests memorization because it's challenging for humans and easy to score objectively. But memorization isn't that important in legal practice; analysis is. LLMs are superhuman at memorization but terrible at analysis.

Eh, also in legal practice there are key skills like selecting the best billable clients, covering your ass, building a reputation, choosing the right market segment, etc. which I’d also argue LLMs suck at.

Re: Re-Evaluating GPT-4's Bar Exam Performance

#10
post #6

The bigger issue here is that actual legal practice looks nothing like the bar, so whether or not an llm passes says nothing about how llms will impact the legal field. Passing the bar should not be understood to mean "can successfully perform legal tasks."

Indeed, and this is also the general problem with most current ways to evaluate AI: by every test there's at least one model which looks wildly superhuman, but actually using them reveals they're book-smart at everything without having any street-smarts.

The difference between expectation and reality is tripping people up in both directions — a nearly-free everything-intern is still very useful, but to treat LLMs* as experts (or capable of meaningful on-the-job learning if you're not fine-tuning the model) is a mistake.

* special purpose AI like Stockfish, however, should be treated as experts

Post reply on HN