Re-Evaluating GPT-4's Bar Exam Performance
link.springer.com
Re-Evaluating GPT-4's Bar Exam Performance
1–10 of 139 posts
Re: Re-Evaluating GPT-4's Bar Exam Performance
#2That still does put it into bar-passing territory, though, since it still scores better than about one sixth of the people that passed the exam.
Re: Re-Evaluating GPT-4's Bar Exam Performance
#3Really glad to see research replicated like this. I’m not surprised that the 90th percentile doesn’t hold up.
It’s still handy though.
Re: Re-Evaluating GPT-4's Bar Exam Performance
#4Very interesting. The abstract claims that although GPT-4 was claimed to score in the 92nd percentile on the bar exam, when correcting for a bunch of things they find that these results are overinflated, and that it only scores in the 15th percentile specifically on essays when compared to only people that passed the bar. That still does put it into bar-passing territory, though, since it still scores better than abo…
Re: Re-Evaluating GPT-4's Bar Exam Performance
#5For example, I doubt that it asks whether, for a person of average wealth and income, a $1000 fine is a more or less severe punishment than a month in jail.
Re: Re-Evaluating GPT-4's Bar Exam Performance
#6Passing the bar should not be understood to mean "can successfully perform legal tasks."
Re: Re-Evaluating GPT-4's Bar Exam Performance
#7So, GPT-4 scores closer to the bottom of people who pass the bar the first time. In other words, it matches the people who cull the rules from texts already written, but who cannot apply it imaginatively.
Re: Re-Evaluating GPT-4's Bar Exam Performance
#8A basic problem with evaluations like these is that the test is designed to discriminate between humans who would make good lawyers and humans who would not make good lawyers. The test is not necessarily any good at telling whether a non-human would make a good lawyer, since it will not test anything that pretty much all humans know, but non-humans may not. For example, I doubt that it asks whether, for a person of a…
Re: Re-Evaluating GPT-4's Bar Exam Performance
#9A basic problem with evaluations like these is that the test is designed to discriminate between humans who would make good lawyers and humans who would not make good lawyers. The test is not necessarily any good at telling whether a non-human would make a good lawyer, since it will not test anything that pretty much all humans know, but non-humans may not. For example, I doubt that it asks whether, for a person of a…
Honestly, this is giving the bar exam (and GPT-4) too much credit. The bar tests memorization because it's challenging for humans and easy to score objectively. But memorization isn't that important in legal practice; analysis is. LLMs are superhuman at memorization but terrible at analysis.
Re: Re-Evaluating GPT-4's Bar Exam Performance
#10The bigger issue here is that actual legal practice looks nothing like the bar, so whether or not an llm passes says nothing about how llms will impact the legal field. Passing the bar should not be understood to mean "can successfully perform legal tasks."
The difference between expectation and reality is tripping people up in both directions — a nearly-free everything-intern is still very useful, but to treat LLMs* as experts (or capable of meaningful on-the-job learning if you're not fine-tuning the model) is a mistake.
* special purpose AI like Stockfish, however, should be treated as experts