Live data from Hacker News

Re-Evaluating GPT-4's Bar Exam Performance

link.springer.com

31–40 of 139 posts

Re: Re-Evaluating GPT-4's Bar Exam Performance

#31

Earlier quoted context omitted.

You can just ask it, you know. GPT-4o: “Average wealth and income” can vary significantly by region and context. However, in the United States, as a rough benchmark, the median household income is around $70,000 per year. Wealth, which includes assets such as savings, property, and investments minus debts, is harder to pinpoint but median net worth for U.S. households is approximately $100,000. These figures provide…

I like that it immediately assumed the US, even though nothing in your question suggested it. I love that all LLMs have a strong US centric bias. Btw I'm not personally a lawyer, but I've heard that GPT is especially prone to mixing laws across the borders - for example you ask a law question in language X, and get a response that uses a law from a country Y - and it's extremally convincing doing that (unless you're…

ChatGPT has user-customizable "instructions", and mine are set to tell it where I live. Any user can do the same, so that it will not make incorrect assumptions for you.

Re: Re-Evaluating GPT-4's Bar Exam Performance

#32
It is difficult to comment without sounding obnoxious, but having taken the bar exam, I found the exam simple. Surprisingly simple. I think it was the single most over hyped experience of my life. I was fed all this insecurity and walked into the convention center expecting to participate in the biggest intellectual challenge in my life. Instead, it was endless multiple choice questions and a couple contrived scenarios for essays.

It may also be surprising to some to understand that legal writing is prized for its degree of formalism. It aims to remove all connotation from a message so as to minimize misunderstanding, much like clean code.

It may also be surprising, but the goal when writing a legal brief or judicial opinion is not to try to sound smart. The goal is to be clear, objective, and thereby, persuasive. Using big words for the sake of using big words, using rare words, using weasel words like "kind of" or "most of the time" or "many people are saying", writing poetically, being overly obtuse and abstract, these are things that get your law school application rejected, your brief ridiculed, and your bar exam failed.

The simpler your communication, the more formulaic, the better. The more your argument is structured, akin to a computer program, the better.

As compared to some other domain, such as fiction, good legal writing much easier for an attention model to simulate. The best exam answers are the ones that are the most formulaic and that use the smallest lexicon and that use words correctly.

I only want to add this comment because I want to inform how non-lawyers perceive the bar exam. Getting an attention model to pass the bar exam is a low bar. It is not some great technical feat. A programmer can practically write a semantic disambiguation algorithm for legal writing from scratch with moderate effort.

It will be a good accomplishment, but it will only be a stepping stone. I am still waiting for AI to tackle messages that have greater nuance and that are truly free form. LLMs are still not there yet.

Re: Re-Evaluating GPT-4's Bar Exam Performance

#33
post #12

Earlier quoted context omitted.

Honestly, this is giving the bar exam (and GPT-4) too much credit. The bar tests memorization because it's challenging for humans and easy to score objectively. But memorization isn't that important in legal practice; analysis is. LLMs are superhuman at memorization but terrible at analysis.

I've always drawn the link between skill in memorization and in analysis as: - Memorization requires you to retain the details of a large amount of material - The most time-efficient analysis uses instant-recall of relevant general themes to guide research - Ergo, if someone can memorize and recall a large number of details, they can probably also recall relevant general themes, and therefore quickly perform quality…

Problem is the LLM memorized the countless examples you can find of old BAR questions using extreme amounts of compute at training time, they don't have that ability to digest a specific case due to both lack of data and it doesn't retrain for new questions.

A human that can digest the general law can also digest a special case, but that isn't true for an LLM.

Re: Re-Evaluating GPT-4's Bar Exam Performance

#34

It is difficult to comment without sounding obnoxious, but having taken the bar exam, I found the exam simple. Surprisingly simple. I think it was the single most over hyped experience of my life. I was fed all this insecurity and walked into the convention center expecting to participate in the biggest intellectual challenge in my life. Instead, it was endless multiple choice questions and a couple contrived scenari…

> It may also be surprising to some to understand that legal writing is prized for its degree of formalism. It aims to remove all connotation from a message so as to minimize misunderstanding, much like clean code.

> The more your argument is structured, akin to a computer program, the better.

You certainly make legal writing sound like a flavor of technical writing. Simplicity, clarity, structure. Is this an accurate comparison ?

Re: Re-Evaluating GPT-4's Bar Exam Performance

#35

Earlier quoted context omitted.

For a person of average wealth and income, is a $1000 fine is a more or less severe punishment than a month in jail? Be brief. "For a person of average wealth and income, a $1000 fine is generally less severe than a month in jail. A month in jail entails loss of freedom, potential loss of employment, and social stigma, while a $1000 fine, though financially burdensome, does not affect one's freedom or ability to work…

"potential loss of employment," Where is that coming from ? That's a very lawyery way to phrase things. "potential ?" where I live I think people may max out their holidays and overtime (if lucky enough) and leave-without-pay but there would be a conversation with your employer to justify it and how to handle the workload. In the USA, from what I read, it's more than likely that you would just be fired on the spot, r…

Many jails have work release. They get you up at 6am, check you out of the jail, let you go to work, then expect you to check back into jail by 6pm.

Re: Re-Evaluating GPT-4's Bar Exam Performance

#36
post #29

Earlier quoted context omitted.

> which I’d also argue LLMs suck at OK, I’ll bite. What’s your evidence for this argument?

Every bit of interaction I’ve ever had with an LLM. And all the research I’ve seen. They’re plausible word sequence generators, not ‘planning for the future’ agents. Or market analyzers. Or character evaluators. Or anything else. And they tend to be really ‘gullible’. What evidence do you have they could do any of those things? (And not just generate plausible text at a prompt, but actually do those things)

> What evidence do you have they could do any of those things?

Every bit of interaction I’ve ever had with an LLM.

Re: Re-Evaluating GPT-4's Bar Exam Performance

#37

Earlier quoted context omitted.

"potential loss of employment," Where is that coming from ? That's a very lawyery way to phrase things. "potential ?" where I live I think people may max out their holidays and overtime (if lucky enough) and leave-without-pay but there would be a conversation with your employer to justify it and how to handle the workload. In the USA, from what I read, it's more than likely that you would just be fired on the spot, r…

Many jails have work release. They get you up at 6am, check you out of the jail, let you go to work, then expect you to check back into jail by 6pm.

Oh, that's really great ! Is that in the US ?

Re: Re-Evaluating GPT-4's Bar Exam Performance

#38

Earlier quoted context omitted.

"potential loss of employment," Where is that coming from ? That's a very lawyery way to phrase things. "potential ?" where I live I think people may max out their holidays and overtime (if lucky enough) and leave-without-pay but there would be a conversation with your employer to justify it and how to handle the workload. In the USA, from what I read, it's more than likely that you would just be fired on the spot, r…

Many jails have work release. They get you up at 6am, check you out of the jail, let you go to work, then expect you to check back into jail by 6pm.

Which for many professional (and other jobs) probably would require a bunch of tap-dancing around your strict schedule if you were hiding the actual reason.

Re: Re-Evaluating GPT-4's Bar Exam Performance

#39

It is difficult to comment without sounding obnoxious, but having taken the bar exam, I found the exam simple. Surprisingly simple. I think it was the single most over hyped experience of my life. I was fed all this insecurity and walked into the convention center expecting to participate in the biggest intellectual challenge in my life. Instead, it was endless multiple choice questions and a couple contrived scenari…

> It may also be surprising to some to understand that legal writing is prized for its degree of formalism. It aims to remove all connotation from a message so as to minimize misunderstanding, much like clean code. > The more your argument is structured, akin to a computer program, the better. You certainly make legal writing sound like a flavor of technical writing. Simplicity, clarity, structure. Is this an accurat…

it is called a legal code after all

Re: Re-Evaluating GPT-4's Bar Exam Performance

#40
post #6

The bigger issue here is that actual legal practice looks nothing like the bar, so whether or not an llm passes says nothing about how llms will impact the legal field. Passing the bar should not be understood to mean "can successfully perform legal tasks."

This, along with several other "meta" objections, is a significant portion of the discussion in the paper.

They basically say two things. First, although the measurement is repeatable at face value, there are several factors that make it less impressive than assumed, and the model performs fairly poorly compared to likely prospective lawyers. Second, there is a number of reasons why the percentile on the test doesn't measure lawyering skills.

One of the other interesting points they bring up is that there is no incentive for humans to seek scores much above passing on the test, because your career outlook doesn't depend on it in any way. This is different from many other placement exams.

Post reply on HN