Live data from Hacker News

Re-Evaluating GPT-4's Bar Exam Performance

link.springer.com

71–80 of 139 posts

Re: Re-Evaluating GPT-4's Bar Exam Performance

#71

It is difficult to comment without sounding obnoxious, but having taken the bar exam, I found the exam simple. Surprisingly simple. I think it was the single most over hyped experience of my life. I was fed all this insecurity and walked into the convention center expecting to participate in the biggest intellectual challenge in my life. Instead, it was endless multiple choice questions and a couple contrived scenari…

> It may also be surprising to some to understand that legal writing is prized for its degree of formalism. It aims to remove all connotation from a message so as to minimize misunderstanding, much like clean code. > The more your argument is structured, akin to a computer program, the better. You certainly make legal writing sound like a flavor of technical writing. Simplicity, clarity, structure. Is this an accurat…

IANAL but have about a decade of experience negotiating contracts in M&A, and I think the comparison is very apt for that particular context, at a minimum. Maybe more so than other parts of Law where there can be an element of persuasion to any given argument.

Re: Re-Evaluating GPT-4's Bar Exam Performance

#72

It is difficult to comment without sounding obnoxious, but having taken the bar exam, I found the exam simple. Surprisingly simple. I think it was the single most over hyped experience of my life. I was fed all this insecurity and walked into the convention center expecting to participate in the biggest intellectual challenge in my life. Instead, it was endless multiple choice questions and a couple contrived scenari…

Genuinely asking: you think the bar exam is a low bar because you personally found it easy, even though the vast majority of takers do not? Doesn't this just reflect your own inability to empathize with other people?

Re: Re-Evaluating GPT-4's Bar Exam Performance

#73
post #54

Earlier quoted context omitted.

On any topic that I understand well, LLM output is garbage: it requires more energy to fix it than to solve the original problem to begin with. Are we sure these exams are not present in the training data? (ability to recall information is not impressive for a computer) Still I'm terrible at many many tasks e.g., drawing from description and the models widen significantly types of problems that I can even try (where…

the models that you have tried .. are garbage. hmmm Maybe you are not among the many, many, many inside professionals and unofrmed services that have different access than you? money talks?

It is remarkable that folks who tried a garbage LLM like copilot, 3.5, Gemini, or made meta LLMs say naughty words, seem to think these are still SOA. Sometimes I stumble on them and I am shocked at the degradation in quality then realize my settings are wrong. People are vastly underestimating the rate of change here.

Re: Re-Evaluating GPT-4's Bar Exam Performance

#74
post #12

Earlier quoted context omitted.

Honestly, this is giving the bar exam (and GPT-4) too much credit. The bar tests memorization because it's challenging for humans and easy to score objectively. But memorization isn't that important in legal practice; analysis is. LLMs are superhuman at memorization but terrible at analysis.

I've always drawn the link between skill in memorization and in analysis as: - Memorization requires you to retain the details of a large amount of material - The most time-efficient analysis uses instant-recall of relevant general themes to guide research - Ergo, if someone can memorize and recall a large number of details, they can probably also recall relevant general themes, and therefore quickly perform quality…

I’m not sure why you’re being downvoted for this. I agree with you, fact recall is useful and necessary. If you have a larger and more tightly connected base of facts in your head, you can draw better connections.

And even though legal practice tends to be fairly slow and deliberative, there are settings (such as trial advocacy) where there is a real advantage to being able to cite a case or statute from memory.

All that said, I still maintain that it’s a poor way to compare humans with machines, for the same reason it would be poor to compare GPT-4 to a novelist on their tokens per second written.

Re: Re-Evaluating GPT-4's Bar Exam Performance

#75
post #9

Earlier quoted context omitted.

Eh, also in legal practice there are key skills like selecting the best billable clients, covering your ass, building a reputation, choosing the right market segment, etc. which I’d also argue LLMs suck at.

I don’t know. There was some talk this weekend about CEOs being replaced by AI. Given the overlap in skill, I’d say there is a distinct possibility an LLM could do that. https://www.msn.com/en-us/money/companies/ceos-could-easily-...

Phoebe Moore who that quote was attributed to has never been a CEO or even worked at a non-academic organisation.

So much of a what a CEO does is fostering culture, hiring people and setting a unique vision for the company.

Imagine thinking people would be inspired to work for a chatbot. Hilariously ridiculous.

Re: Re-Evaluating GPT-4's Bar Exam Performance

#76
post #54

Scoring 96 percentile among humans taking the exam without moving goal posts would have been science fiction two years ago. Now it’s suddenly not good enough and the fact a computer program can score decent among passing lawyers and first time test takers is something to sneer at. The fact I can talk to the computer and it responds to me idiomatically and understands my semantic intent well enough to be nearly indist…

On any topic that I understand well, LLM output is garbage: it requires more energy to fix it than to solve the original problem to begin with. Are we sure these exams are not present in the training data? (ability to recall information is not impressive for a computer) Still I'm terrible at many many tasks e.g., drawing from description and the models widen significantly types of problems that I can even try (where…

I suspect by garbage you mean not perfect.

To be more precise can you please give a topic you know well and your % guess how often the answers are wrong on the topic?

Re: Re-Evaluating GPT-4's Bar Exam Performance

#77

Scoring 96 percentile among humans taking the exam without moving goal posts would have been science fiction two years ago. Now it’s suddenly not good enough and the fact a computer program can score decent among passing lawyers and first time test takers is something to sneer at. The fact I can talk to the computer and it responds to me idiomatically and understands my semantic intent well enough to be nearly indist…

By 96th percentile do you mea 69th? From the abstract:

> data from a recent July administration of the same exam suggests GPT-4’s overall UBE percentile was below the 69th percentile, and 48th percentile on essays. Third, examining official NCBE data and using several conservative statistical assumptions, GPT-4’s performance against first-time test takers is estimated to be 62nd percentile, including 42nd percentile on essays. Fourth, when examining only those who passed the exam (i.e. licensed or license-pending attorneys), GPT-4’s performance is estimated to drop to 48th percentile overall, and 15th percentile on essays.

Re: Re-Evaluating GPT-4's Bar Exam Performance

#78

Scoring 96 percentile among humans taking the exam without moving goal posts would have been science fiction two years ago. Now it’s suddenly not good enough and the fact a computer program can score decent among passing lawyers and first time test takers is something to sneer at. The fact I can talk to the computer and it responds to me idiomatically and understands my semantic intent well enough to be nearly indist…

It’s the hype. We could invent warp drive but if it was hyped as the cure for cancer, poverty, war, and the gateway to untold riches and immortality while simultaneously being the most dangerous invention in history destined to completely destroy humanity people would be “oh ho hum we made it to Centauri in a week” pretty fast.

Add some obnoxious pseudo-intellectual windbags building a cult around it and people would be down right turned off.

Hype is also taken as a strong contrarian indicator by most scientific and engineering types. A lot of hype means it’s snake oil. This heuristic is actually correct more often than it’s not, but it is occasionally wrong.

Re: Re-Evaluating GPT-4's Bar Exam Performance

#79

Scoring 96 percentile among humans taking the exam without moving goal posts would have been science fiction two years ago. Now it’s suddenly not good enough and the fact a computer program can score decent among passing lawyers and first time test takers is something to sneer at. The fact I can talk to the computer and it responds to me idiomatically and understands my semantic intent well enough to be nearly indist…

I have difficulty being optimistic about LLMs because they don’t benefit my work now, and I don’t see a way that they enhance our humanity. They’re explicitly pitched as something that should eat all sorts of jobs.

The problem isn’t the LLMs per se, it’s what we want to do with them. And, being human, it becomes difficult to separate the two.

Also, they seem to attract people who get real aggressive about defending them and seem to attach part of their identity onto them, which is weird.

Re: Re-Evaluating GPT-4's Bar Exam Performance

#80
post #54

Earlier quoted context omitted.

On any topic that I understand well, LLM output is garbage: it requires more energy to fix it than to solve the original problem to begin with. Are we sure these exams are not present in the training data? (ability to recall information is not impressive for a computer) Still I'm terrible at many many tasks e.g., drawing from description and the models widen significantly types of problems that I can even try (where…

> On any topic that I understand well, LLM output is garbage: it requires more energy to fix it than to solve the original problem to begin with. That's probably true, which is why human most knowledge workers aren't going away any time soon. That said, I have better luck with a different approach: I use LLM's to learn things that I don't already understand well. This forces me to actively understand and validate the…

I would classify all of those as "non-traditional" learning techniques, unless you actually mean using a textbook while taking a class with a human teacher.

Well written textbooks are consumable on their own for some people, but most are not written for that.

Post reply on HN