Live data from Hacker News

Re-Evaluating GPT-4's Bar Exam Performance

link.springer.com

41–50 of 139 posts

Re: Re-Evaluating GPT-4's Bar Exam Performance

#41
> Furthermore, unlike its documentation for the other exams it tested (OpenAI 2023b, p. 25), OpenAI’s technical report provides no direct citation for how the UBE percentile was computed, creating further uncertainty over both the original source and validity of the 90th percentile claim.

This is the part that bothered me (licensed attorney) from the start. If it scores this high, where are the receipts? I’m sure OpenAI has the social capital to coordinate with the National Conference of Bar Examiners to have a GPT “sit” for a simulated bar exam.

Re: Re-Evaluating GPT-4's Bar Exam Performance

#42

Earlier quoted context omitted.

You can just ask it, you know. GPT-4o: “Average wealth and income” can vary significantly by region and context. However, in the United States, as a rough benchmark, the median household income is around $70,000 per year. Wealth, which includes assets such as savings, property, and investments minus debts, is harder to pinpoint but median net worth for U.S. households is approximately $100,000. These figures provide…

I like that it immediately assumed the US, even though nothing in your question suggested it. I love that all LLMs have a strong US centric bias. Btw I'm not personally a lawyer, but I've heard that GPT is especially prone to mixing laws across the borders - for example you ask a law question in language X, and get a response that uses a law from a country Y - and it's extremally convincing doing that (unless you're…

This is my experience with Hackernews. If the comment doesn't specify the country, it's an American talking about the USA

Re: Re-Evaluating GPT-4's Bar Exam Performance

#43
Scoring 96 percentile among humans taking the exam without moving goal posts would have been science fiction two years ago. Now it’s suddenly not good enough and the fact a computer program can score decent among passing lawyers and first time test takers is something to sneer at.

The fact I can talk to the computer and it responds to me idiomatically and understands my semantic intent well enough to be nearly indistinguishable from a human being is breath taking. Anyone who views it as anything less in 2024 and asserts with a straight face they wouldn’t have said the same thing in 2020 is lying.

I do however find the paper really useful in contextualizing the scoring with a much finer grain. Personally I didn’t take the 96 percentile score to be anything other than “among the mass who take the test,” and have enough experience with professional licensing exams to know a huge percentage of test takers fail and are repeat test takers. Placing the goal posts quantitatively for the next levels of achievement is a useful exercise. But the profusion of jaded nerds makes me sad.

Re: Re-Evaluating GPT-4's Bar Exam Performance

#44
post #26

It appears that researchers and commentators are totally missing the application of LLMs to law, and to other areas of professional practice. A generic trained-on-Quora LLM is going to be straight garbage for any specialization, but one that is trained on the contents of the law library will be utterly brilliant for assisting a practicing attorney. People pay serious money for legal indexes, cross-references, and res…

It is a lossy compressed index. It has an approximate knowledge of law, and that approximation can be pretty good - but it doesn't know when it's outputting plausible but made-up claims. As with GitHub Copilot, it's probably going to be a mixed bag until we can overcome that, because spotting subtle but grave errors can be harder than writing something from scratch.

There's already a fair number of stories of LLMs used by an attorney messing up court filings - e.g., inventing fake case law.

Re: Re-Evaluating GPT-4's Bar Exam Performance

#45

Earlier quoted context omitted.

I like that it immediately assumed the US, even though nothing in your question suggested it. I love that all LLMs have a strong US centric bias. Btw I'm not personally a lawyer, but I've heard that GPT is especially prone to mixing laws across the borders - for example you ask a law question in language X, and get a response that uses a law from a country Y - and it's extremally convincing doing that (unless you're…

ChatGPT has user-customizable "instructions", and mine are set to tell it where I live. Any user can do the same, so that it will not make incorrect assumptions for you.

You might increase the probability of getting a correct answer for your region, but imo you decrease your awareness to allucination. Overall you can still get a wrong answer

Re: Re-Evaluating GPT-4's Bar Exam Performance

#46
post #7

They originally scored against a test usually taken by people who failed the bar. So, GPT-4 scores closer to the bottom of people who pass the bar the first time. In other words, it matches the people who cull the rules from texts already written, but who cannot apply it imaginatively.

> In other words, it matches the people who cull the rules from texts already written, but who cannot apply it imaginatively.

Where did you find that in the article?

Re: Re-Evaluating GPT-4's Bar Exam Performance

#47
post #26

It appears that researchers and commentators are totally missing the application of LLMs to law, and to other areas of professional practice. A generic trained-on-Quora LLM is going to be straight garbage for any specialization, but one that is trained on the contents of the law library will be utterly brilliant for assisting a practicing attorney. People pay serious money for legal indexes, cross-references, and res…

It is a lossy compressed index. It has an approximate knowledge of law, and that approximation can be pretty good - but it doesn't know when it's outputting plausible but made-up claims. As with GitHub Copilot, it's probably going to be a mixed bag until we can overcome that, because spotting subtle but grave errors can be harder than writing something from scratch. There's already a fair number of stories of LLMs us…

I am not suggesting that the generative aspects would be useful in drafting motions and such. I am suggesting that their tendency towards false results is harmless if you just use them as a complex index. For example, you could ask it to list appellate cases where one party argued such-and-such and prevailed. Then you would go read the cases.

Re: Re-Evaluating GPT-4's Bar Exam Performance

#48

It is difficult to comment without sounding obnoxious, but having taken the bar exam, I found the exam simple. Surprisingly simple. I think it was the single most over hyped experience of my life. I was fed all this insecurity and walked into the convention center expecting to participate in the biggest intellectual challenge in my life. Instead, it was endless multiple choice questions and a couple contrived scenari…

I took a sample CA bar exam for fun, as a non-lawyer who has never set foot in law school. Maybe the sample exam was tougher than the real thing, but I found it surprisingly difficult. A lot of the correct answers to questions were non-obvious -- they weren't based on straightforward logic, nor were they based on moral reasoning, and there was no place for "natural law" -- so to answer questions properly you had to have memorized a bit of coursework. There were also a few questions that seemed almost designed to deceive the test-taker; the "obvious" moral choices were the wrong ones.

So maybe it's easy if you study that stuff for a year or two. But you can't just walk in and expect to pass, or bullshit your way through it.

I agree with you on legal writing, but there appears to be a certain amount of ambiguity inherent to language. The Uniform Commercial Code, for instance, is maddeningly vague at points.

Re: Re-Evaluating GPT-4's Bar Exam Performance

#49
post #14

Earlier quoted context omitted.

I don’t know. There was some talk this weekend about CEOs being replaced by AI. Given the overlap in skill, I’d say there is a distinct possibility an LLM could do that. https://www.msn.com/en-us/money/companies/ceos-could-easily-...

Bwahaha. This is like the ‘everything can be a directed graph db’, ‘everything should be a micro service’, etc. fads. No one who has been a CEO, or frankly even worked closely with one, would think this could be even remotely close to possible. Or desirable if it was. But that is probably 1% or less of the population eh?

https://www.dqindia.com/company-makes-ai-robot-its-ceo-makes...

Seems your claim's been disproven already

Re: Re-Evaluating GPT-4's Bar Exam Performance

#50
post #49
post #14

Earlier quoted context omitted.

Bwahaha. This is like the ‘everything can be a directed graph db’, ‘everything should be a micro service’, etc. fads. No one who has been a CEO, or frankly even worked closely with one, would think this could be even remotely close to possible. Or desirable if it was. But that is probably 1% or less of the population eh?

https://www.dqindia.com/company-makes-ai-robot-its-ceo-makes... Seems your claim's been disproven already

Bwaha. Funny the company named as doing so doesn’t mention it on their actual management team [http://www.netdragon.com/about/management-team.shtml], listing an actual human CEO instead.

But it makes for a fun soundbite eh? Especially when the article claims it was in the past, and totally was awesome. Sucker born every minute.

Post reply on HN