Live data from Hacker News

Re-Evaluating GPT-4's Bar Exam Performance

link.springer.com

101–110 of 139 posts

Re: Re-Evaluating GPT-4's Bar Exam Performance

#102

Earlier quoted context omitted.

It is remarkable that folks who tried a garbage LLM like copilot, 3.5, Gemini, or made meta LLMs say naughty words, seem to think these are still SOA. Sometimes I stumble on them and I am shocked at the degradation in quality then realize my settings are wrong. People are vastly underestimating the rate of change here.

> People are vastly underestimating the rate of change here GPT-3.5 was released in March 2022. We are now in June 2024. Over 2 years later. And on average GPT-4 is about 40% more accurate. For me, LLMs are very much like self-driving cars. On the journey towards perfect accuracy it gets progressively harder to make advancements. And for it to replace the status quo it really does need to be perfect. And there is no…

Its enough to decrease the amount of ppl you need in IT by a factor of 20-30%.

Ppl dont want to hear that, but you see less and less offers and not only for junior positions.

Hard truth is that like with any tool/automation - the higher performance improves, the less ppl are needed for this kind of work.

Just look at how some parts of manual labor were made redundant.

Why ppl think it wont be the same with mental work is beyond me.

Re: Re-Evaluating GPT-4's Bar Exam Performance

#103

Earlier quoted context omitted.

I don’t know. There was some talk this weekend about CEOs being replaced by AI. Given the overlap in skill, I’d say there is a distinct possibility an LLM could do that. https://www.msn.com/en-us/money/companies/ceos-could-easily-...

Phoebe Moore who that quote was attributed to has never been a CEO or even worked at a non-academic organisation. So much of a what a CEO does is fostering culture, hiring people and setting a unique vision for the company. Imagine thinking people would be inspired to work for a chatbot. Hilariously ridiculous.

If that chatbot had Steve Jobs voice ?

I dunno, I would probably prefer to work under that chatbot than my current CEO that only tries to squize as much as possible out of ppl already working for him.

Re: Re-Evaluating GPT-4's Bar Exam Performance

#104
post #54

Scoring 96 percentile among humans taking the exam without moving goal posts would have been science fiction two years ago. Now it’s suddenly not good enough and the fact a computer program can score decent among passing lawyers and first time test takers is something to sneer at. The fact I can talk to the computer and it responds to me idiomatically and understands my semantic intent well enough to be nearly indist…

On any topic that I understand well, LLM output is garbage: it requires more energy to fix it than to solve the original problem to begin with. Are we sure these exams are not present in the training data? (ability to recall information is not impressive for a computer) Still I'm terrible at many many tasks e.g., drawing from description and the models widen significantly types of problems that I can even try (where…

> On any topic that I understand well, LLM output is garbage

I've heard that claim many times, but never is there any specific follow-up on which topics they mean. Of course, there are areas like math and programming where LLMs might not perform as well as a senior programmer or mathematician, sometimes producing programs that do not compile or incorrect calculations/ideas. However, this isn't exactly "garbage" as some suggest. At worst, it's more like a freshman-level answer, and at best, it can be a perfectly valid and correct response.

Re: Re-Evaluating GPT-4's Bar Exam Performance

#105
post #88

Earlier quoted context omitted.

Functional linear analysis - it has tendency to produce a proof for unprovable statements; the proofs will be logically argued and well structured and step 8 will have a statement that is obvious nonsense even to a beginning student, me. The professor on the other hand will ask why I'm trying to prove the false statement and expertly help me find my logic error.

Specifics like this make it much easier to agree on LLM capabilities, thank you. Automatic proof generation is a massive open problem in all of computer science and not close to be solved. It’s true LLMs aren’t great at it and more is required for example as with the geometry system Deepmind progresses on. On the other hand they can be very useful to explain concepts and allow interactive questioning to drill down an…

How do yo debug its hallucination misinformation via voice interface while you commute?

Re: Re-Evaluating GPT-4's Bar Exam Performance

#106
post #54

Earlier quoted context omitted.

On any topic that I understand well, LLM output is garbage: it requires more energy to fix it than to solve the original problem to begin with. Are we sure these exams are not present in the training data? (ability to recall information is not impressive for a computer) Still I'm terrible at many many tasks e.g., drawing from description and the models widen significantly types of problems that I can even try (where…

On what topics you understand well does GOT-4o or Claude Opus produce garbage?

High school math problems.

Re: Re-Evaluating GPT-4's Bar Exam Performance

#107

It is difficult to comment without sounding obnoxious, but having taken the bar exam, I found the exam simple. Surprisingly simple. I think it was the single most over hyped experience of my life. I was fed all this insecurity and walked into the convention center expecting to participate in the biggest intellectual challenge in my life. Instead, it was endless multiple choice questions and a couple contrived scenari…

I took a sample CA bar exam for fun, as a non-lawyer who has never set foot in law school. Maybe the sample exam was tougher than the real thing, but I found it surprisingly difficult. A lot of the correct answers to questions were non-obvious -- they weren't based on straightforward logic, nor were they based on moral reasoning, and there was no place for "natural law" -- so to answer questions properly you had to h…

I took a sample test as well and I believe I did well enough on some sections that I could have barely passed those sections with no background.

A key item which jumped out at me right away is that in addition to the logic, the possible answers would include things which the scenario didn't address. Like, a wrong answer might make an assumption that you couldn't arrive to via the scenario. More tricky were the answers which made assumptions which you knew to be correct (based on a real event,) but still wasn't addressed in the scenario. If you combined these two elements (getting the logic right, and eliminating assumptions which you couldn't make from the scenario) then you could do well on those.

The sections I wouldn't have passed were those which required specific law knowledge. So, some sections were general, while others required knowledge of something like real estate law. I don't remember if these questions were otherwise similar to the ones I could pass.

An LLM is taking this test as essentially an open book.

Re: Re-Evaluating GPT-4's Bar Exam Performance

#108
post #54

Earlier quoted context omitted.

On any topic that I understand well, LLM output is garbage: it requires more energy to fix it than to solve the original problem to begin with. Are we sure these exams are not present in the training data? (ability to recall information is not impressive for a computer) Still I'm terrible at many many tasks e.g., drawing from description and the models widen significantly types of problems that I can even try (where…

> On any topic that I understand well, LLM output is garbage I've heard that claim many times, but never is there any specific follow-up on which topics they mean. Of course, there are areas like math and programming where LLMs might not perform as well as a senior programmer or mathematician, sometimes producing programs that do not compile or incorrect calculations/ideas. However, this isn't exactly "garbage" as so…

> At worst, it's more like a freshman-level answer

That is garbage.

Re: Re-Evaluating GPT-4's Bar Exam Performance

#109

Scoring 96 percentile among humans taking the exam without moving goal posts would have been science fiction two years ago. Now it’s suddenly not good enough and the fact a computer program can score decent among passing lawyers and first time test takers is something to sneer at. The fact I can talk to the computer and it responds to me idiomatically and understands my semantic intent well enough to be nearly indist…

Yeah it’s insane, I am actually scared the llm is like sentient and secretly plotting to kill me. I bet we have like full AGI next year because Elon said so and Sam Altman probably has AGI already internally at Open AI. I am actually selling my house now and going all in Nvidia and just live in my car until we get the AGI

Re: Re-Evaluating GPT-4's Bar Exam Performance

#110

Earlier quoted context omitted.

> On any topic that I understand well, LLM output is garbage I've heard that claim many times, but never is there any specific follow-up on which topics they mean. Of course, there are areas like math and programming where LLMs might not perform as well as a senior programmer or mathematician, sometimes producing programs that do not compile or incorrect calculations/ideas. However, this isn't exactly "garbage" as so…

> At worst, it's more like a freshman-level answer That is garbage.

I hope you don't hold a teaching position at a university then.
Post reply on HN