This is the part that bothered me (licensed attorney) from the start. If it scores this high, where are the receipts? I’m sure OpenAI has the social capital to coordinate with the National Conference of Bar Examiners to have a GPT “sit” for a simulated bar exam.
Re-Evaluating GPT-4's Bar Exam Performance
41–50 of 139 posts
Re: Re-Evaluating GPT-4's Bar Exam Performance
#42Earlier quoted context omitted.
You can just ask it, you know. GPT-4o: “Average wealth and income” can vary significantly by region and context. However, in the United States, as a rough benchmark, the median household income is around $70,000 per year. Wealth, which includes assets such as savings, property, and investments minus debts, is harder to pinpoint but median net worth for U.S. households is approximately $100,000. These figures provide…
I like that it immediately assumed the US, even though nothing in your question suggested it. I love that all LLMs have a strong US centric bias. Btw I'm not personally a lawyer, but I've heard that GPT is especially prone to mixing laws across the borders - for example you ask a law question in language X, and get a response that uses a law from a country Y - and it's extremally convincing doing that (unless you're…
Re: Re-Evaluating GPT-4's Bar Exam Performance
#43The fact I can talk to the computer and it responds to me idiomatically and understands my semantic intent well enough to be nearly indistinguishable from a human being is breath taking. Anyone who views it as anything less in 2024 and asserts with a straight face they wouldn’t have said the same thing in 2020 is lying.
I do however find the paper really useful in contextualizing the scoring with a much finer grain. Personally I didn’t take the 96 percentile score to be anything other than “among the mass who take the test,” and have enough experience with professional licensing exams to know a huge percentage of test takers fail and are repeat test takers. Placing the goal posts quantitatively for the next levels of achievement is a useful exercise. But the profusion of jaded nerds makes me sad.
Re: Re-Evaluating GPT-4's Bar Exam Performance
#44It appears that researchers and commentators are totally missing the application of LLMs to law, and to other areas of professional practice. A generic trained-on-Quora LLM is going to be straight garbage for any specialization, but one that is trained on the contents of the law library will be utterly brilliant for assisting a practicing attorney. People pay serious money for legal indexes, cross-references, and res…
There's already a fair number of stories of LLMs used by an attorney messing up court filings - e.g., inventing fake case law.
Re: Re-Evaluating GPT-4's Bar Exam Performance
#45Earlier quoted context omitted.
I like that it immediately assumed the US, even though nothing in your question suggested it. I love that all LLMs have a strong US centric bias. Btw I'm not personally a lawyer, but I've heard that GPT is especially prone to mixing laws across the borders - for example you ask a law question in language X, and get a response that uses a law from a country Y - and it's extremally convincing doing that (unless you're…
ChatGPT has user-customizable "instructions", and mine are set to tell it where I live. Any user can do the same, so that it will not make incorrect assumptions for you.
Re: Re-Evaluating GPT-4's Bar Exam Performance
#46They originally scored against a test usually taken by people who failed the bar. So, GPT-4 scores closer to the bottom of people who pass the bar the first time. In other words, it matches the people who cull the rules from texts already written, but who cannot apply it imaginatively.
Where did you find that in the article?
Re: Re-Evaluating GPT-4's Bar Exam Performance
#47It appears that researchers and commentators are totally missing the application of LLMs to law, and to other areas of professional practice. A generic trained-on-Quora LLM is going to be straight garbage for any specialization, but one that is trained on the contents of the law library will be utterly brilliant for assisting a practicing attorney. People pay serious money for legal indexes, cross-references, and res…
It is a lossy compressed index. It has an approximate knowledge of law, and that approximation can be pretty good - but it doesn't know when it's outputting plausible but made-up claims. As with GitHub Copilot, it's probably going to be a mixed bag until we can overcome that, because spotting subtle but grave errors can be harder than writing something from scratch. There's already a fair number of stories of LLMs us…
Re: Re-Evaluating GPT-4's Bar Exam Performance
#48It is difficult to comment without sounding obnoxious, but having taken the bar exam, I found the exam simple. Surprisingly simple. I think it was the single most over hyped experience of my life. I was fed all this insecurity and walked into the convention center expecting to participate in the biggest intellectual challenge in my life. Instead, it was endless multiple choice questions and a couple contrived scenari…
So maybe it's easy if you study that stuff for a year or two. But you can't just walk in and expect to pass, or bullshit your way through it.
I agree with you on legal writing, but there appears to be a certain amount of ambiguity inherent to language. The Uniform Commercial Code, for instance, is maddeningly vague at points.
Re: Re-Evaluating GPT-4's Bar Exam Performance
#49Earlier quoted context omitted.
I don’t know. There was some talk this weekend about CEOs being replaced by AI. Given the overlap in skill, I’d say there is a distinct possibility an LLM could do that. https://www.msn.com/en-us/money/companies/ceos-could-easily-...
Bwahaha. This is like the ‘everything can be a directed graph db’, ‘everything should be a micro service’, etc. fads. No one who has been a CEO, or frankly even worked closely with one, would think this could be even remotely close to possible. Or desirable if it was. But that is probably 1% or less of the population eh?
Seems your claim's been disproven already
Re: Re-Evaluating GPT-4's Bar Exam Performance
#50Earlier quoted context omitted.
Bwahaha. This is like the ‘everything can be a directed graph db’, ‘everything should be a micro service’, etc. fads. No one who has been a CEO, or frankly even worked closely with one, would think this could be even remotely close to possible. Or desirable if it was. But that is probably 1% or less of the population eh?
https://www.dqindia.com/company-makes-ai-robot-its-ceo-makes... Seems your claim's been disproven already
But it makes for a fun soundbite eh? Especially when the article claims it was in the past, and totally was awesome. Sucker born every minute.