Re-Evaluating GPT-4's Bar Exam Performance
101–110 of 139 posts
Re: Re-Evaluating GPT-4's Bar Exam Performance
#102Earlier quoted context omitted.
It is remarkable that folks who tried a garbage LLM like copilot, 3.5, Gemini, or made meta LLMs say naughty words, seem to think these are still SOA. Sometimes I stumble on them and I am shocked at the degradation in quality then realize my settings are wrong. People are vastly underestimating the rate of change here.
> People are vastly underestimating the rate of change here GPT-3.5 was released in March 2022. We are now in June 2024. Over 2 years later. And on average GPT-4 is about 40% more accurate. For me, LLMs are very much like self-driving cars. On the journey towards perfect accuracy it gets progressively harder to make advancements. And for it to replace the status quo it really does need to be perfect. And there is no…
Ppl dont want to hear that, but you see less and less offers and not only for junior positions.
Hard truth is that like with any tool/automation - the higher performance improves, the less ppl are needed for this kind of work.
Just look at how some parts of manual labor were made redundant.
Why ppl think it wont be the same with mental work is beyond me.
Re: Re-Evaluating GPT-4's Bar Exam Performance
#103Earlier quoted context omitted.
I don’t know. There was some talk this weekend about CEOs being replaced by AI. Given the overlap in skill, I’d say there is a distinct possibility an LLM could do that. https://www.msn.com/en-us/money/companies/ceos-could-easily-...
Phoebe Moore who that quote was attributed to has never been a CEO or even worked at a non-academic organisation. So much of a what a CEO does is fostering culture, hiring people and setting a unique vision for the company. Imagine thinking people would be inspired to work for a chatbot. Hilariously ridiculous.
I dunno, I would probably prefer to work under that chatbot than my current CEO that only tries to squize as much as possible out of ppl already working for him.
Re: Re-Evaluating GPT-4's Bar Exam Performance
#104Scoring 96 percentile among humans taking the exam without moving goal posts would have been science fiction two years ago. Now it’s suddenly not good enough and the fact a computer program can score decent among passing lawyers and first time test takers is something to sneer at. The fact I can talk to the computer and it responds to me idiomatically and understands my semantic intent well enough to be nearly indist…
On any topic that I understand well, LLM output is garbage: it requires more energy to fix it than to solve the original problem to begin with. Are we sure these exams are not present in the training data? (ability to recall information is not impressive for a computer) Still I'm terrible at many many tasks e.g., drawing from description and the models widen significantly types of problems that I can even try (where…
I've heard that claim many times, but never is there any specific follow-up on which topics they mean. Of course, there are areas like math and programming where LLMs might not perform as well as a senior programmer or mathematician, sometimes producing programs that do not compile or incorrect calculations/ideas. However, this isn't exactly "garbage" as some suggest. At worst, it's more like a freshman-level answer, and at best, it can be a perfectly valid and correct response.
Re: Re-Evaluating GPT-4's Bar Exam Performance
#105Earlier quoted context omitted.
Functional linear analysis - it has tendency to produce a proof for unprovable statements; the proofs will be logically argued and well structured and step 8 will have a statement that is obvious nonsense even to a beginning student, me. The professor on the other hand will ask why I'm trying to prove the false statement and expertly help me find my logic error.
Specifics like this make it much easier to agree on LLM capabilities, thank you. Automatic proof generation is a massive open problem in all of computer science and not close to be solved. It’s true LLMs aren’t great at it and more is required for example as with the geometry system Deepmind progresses on. On the other hand they can be very useful to explain concepts and allow interactive questioning to drill down an…
Re: Re-Evaluating GPT-4's Bar Exam Performance
#106Earlier quoted context omitted.
On any topic that I understand well, LLM output is garbage: it requires more energy to fix it than to solve the original problem to begin with. Are we sure these exams are not present in the training data? (ability to recall information is not impressive for a computer) Still I'm terrible at many many tasks e.g., drawing from description and the models widen significantly types of problems that I can even try (where…
On what topics you understand well does GOT-4o or Claude Opus produce garbage?
Re: Re-Evaluating GPT-4's Bar Exam Performance
#107It is difficult to comment without sounding obnoxious, but having taken the bar exam, I found the exam simple. Surprisingly simple. I think it was the single most over hyped experience of my life. I was fed all this insecurity and walked into the convention center expecting to participate in the biggest intellectual challenge in my life. Instead, it was endless multiple choice questions and a couple contrived scenari…
I took a sample CA bar exam for fun, as a non-lawyer who has never set foot in law school. Maybe the sample exam was tougher than the real thing, but I found it surprisingly difficult. A lot of the correct answers to questions were non-obvious -- they weren't based on straightforward logic, nor were they based on moral reasoning, and there was no place for "natural law" -- so to answer questions properly you had to h…
A key item which jumped out at me right away is that in addition to the logic, the possible answers would include things which the scenario didn't address. Like, a wrong answer might make an assumption that you couldn't arrive to via the scenario. More tricky were the answers which made assumptions which you knew to be correct (based on a real event,) but still wasn't addressed in the scenario. If you combined these two elements (getting the logic right, and eliminating assumptions which you couldn't make from the scenario) then you could do well on those.
The sections I wouldn't have passed were those which required specific law knowledge. So, some sections were general, while others required knowledge of something like real estate law. I don't remember if these questions were otherwise similar to the ones I could pass.
An LLM is taking this test as essentially an open book.
Re: Re-Evaluating GPT-4's Bar Exam Performance
#108Earlier quoted context omitted.
On any topic that I understand well, LLM output is garbage: it requires more energy to fix it than to solve the original problem to begin with. Are we sure these exams are not present in the training data? (ability to recall information is not impressive for a computer) Still I'm terrible at many many tasks e.g., drawing from description and the models widen significantly types of problems that I can even try (where…
> On any topic that I understand well, LLM output is garbage I've heard that claim many times, but never is there any specific follow-up on which topics they mean. Of course, there are areas like math and programming where LLMs might not perform as well as a senior programmer or mathematician, sometimes producing programs that do not compile or incorrect calculations/ideas. However, this isn't exactly "garbage" as so…
That is garbage.
Re: Re-Evaluating GPT-4's Bar Exam Performance
#109Scoring 96 percentile among humans taking the exam without moving goal posts would have been science fiction two years ago. Now it’s suddenly not good enough and the fact a computer program can score decent among passing lawyers and first time test takers is something to sneer at. The fact I can talk to the computer and it responds to me idiomatically and understands my semantic intent well enough to be nearly indist…
Re: Re-Evaluating GPT-4's Bar Exam Performance
#110Earlier quoted context omitted.
> On any topic that I understand well, LLM output is garbage I've heard that claim many times, but never is there any specific follow-up on which topics they mean. Of course, there are areas like math and programming where LLMs might not perform as well as a senior programmer or mathematician, sometimes producing programs that do not compile or incorrect calculations/ideas. However, this isn't exactly "garbage" as so…
> At worst, it's more like a freshman-level answer That is garbage.