Earlier quoted context omitted.
I suspect by garbage you mean not perfect. To be more precise can you please give a topic you know well and your % guess how often the answers are wrong on the topic?
I would take their meaning as 'contains enough errors to not be useful', which doesn't need a very high percentage of wrong answers.
Re-Evaluating GPT-4's Bar Exam Performance
91–100 of 139 posts
Re: Re-Evaluating GPT-4's Bar Exam Performance
#92Earlier quoted context omitted.
I suspect by garbage you mean not perfect. To be more precise can you please give a topic you know well and your % guess how often the answers are wrong on the topic?
Functional linear analysis - it has tendency to produce a proof for unprovable statements; the proofs will be logically argued and well structured and step 8 will have a statement that is obvious nonsense even to a beginning student, me. The professor on the other hand will ask why I'm trying to prove the false statement and expertly help me find my logic error.
Automatic proof generation is a massive open problem in all of computer science and not close to be solved. It’s true LLMs aren’t great at it and more is required for example as with the geometry system Deepmind progresses on.
On the other hand they can be very useful to explain concepts and allow interactive questioning to drill down and help build understanding of complex mathematical concepts, all during a morning commute via the voice interface.
Re: Re-Evaluating GPT-4's Bar Exam Performance
#93Earlier quoted context omitted.
I like that it immediately assumed the US, even though nothing in your question suggested it. I love that all LLMs have a strong US centric bias. Btw I'm not personally a lawyer, but I've heard that GPT is especially prone to mixing laws across the borders - for example you ask a law question in language X, and get a response that uses a law from a country Y - and it's extremally convincing doing that (unless you're…
I mean, to be fair, if you're speaking English to it, the most likely possibility is that you're inside the US: https://en.wikipedia.org/wiki/List_of_countries_by_English-s... I know there's a lot of complaints about things being US-centric, but the US is a very large country.
Indeed - the US is a very large country, and consists of over 50 different jurisdictions, each with their own slightly different laws. An answer to a legal question which is correct in one state will often be subtly incorrect in another, and completely wrong in yet another.
Re: Re-Evaluating GPT-4's Bar Exam Performance
#94Earlier quoted context omitted.
Obviously you need subject knowledge, that should be implicit? Keep in mind even today[1] ( in California and few other states) you don't need to go law school to write the Bar exam and practice law, various forms of apprenticeship under a judge or lawyer are allowed You also don't need to write the exam to practice many aspects of the legal profession. The exam is never meant to be a high bar of quality or selection…
> Obviously you need subject knowledge, that should be implicit? Well, in a lot of the so-called soft sciences, you can easily beat a test without subject knowledge. I had figured that the bar exam might be something like that -- but it's more akin to something like biology, where there are a lot of arcane and counterintuitive little rules that have emerged over time. And you need to know those , or you're sunk. You…
Laws are not logical constructs, they are political constructs, why expect logic from them.
Laws are passed or repealed because it is popular or politically advantageous to do so, not necessarily because it is moral or common sense.
On top of that politicians pass stupid legislation without understanding what they are doing all the time, the infamous Indiana pi bill is a simple example, it almost became law and was stopped by sheer luck, a mathematics professor attending that day on a unrelated matter.
Laws are conflicting, confusing, ambiguous and misleading most of the time, that is expected, the legislating them is a messy process. The third arm of any government judiciary sole purpose is to handle this mess.
---
P.S. I cannot say whether continental legal systems are more robust, but perhaps healthier democracy results in better laws.
All democracies are flawed, U.S. democracy is not particularly healthy, it is not say an proportional multi-party representative system with fair distribution of power amongst all citizens.
From the founding it has been a series of compromises, the history is littered with representation fights such as for suffrage, Jim crow and voting rights, slavery, electoral college, number of states or the filibuster or anti Chinese laws and so on and on.
Don't get me wrong last 250 years have been incredible progress and great leaders put their lives down to make it better, hopefully it will be even better in the future, the struggles do show in the laws it is able to pass, repeal or update.
Re: Re-Evaluating GPT-4's Bar Exam Performance
#95Earlier quoted context omitted.
On any topic that I understand well, LLM output is garbage: it requires more energy to fix it than to solve the original problem to begin with. Are we sure these exams are not present in the training data? (ability to recall information is not impressive for a computer) Still I'm terrible at many many tasks e.g., drawing from description and the models widen significantly types of problems that I can even try (where…
I suspect by garbage you mean not perfect. To be more precise can you please give a topic you know well and your % guess how often the answers are wrong on the topic?
Yesterday I was looking for some help on an issue with the unshare command; it repeatedly made bad assumptions about the nature of the error even I provided it with the full error message and one could already guess the initial cause by looking at that.
I guess such errors can be frighteningly common once you get outside of typical web development.
Re: Re-Evaluating GPT-4's Bar Exam Performance
#96Earlier quoted context omitted.
"potential loss of employment," Where is that coming from ? That's a very lawyery way to phrase things. "potential ?" where I live I think people may max out their holidays and overtime (if lucky enough) and leave-without-pay but there would be a conversation with your employer to justify it and how to handle the workload. In the USA, from what I read, it's more than likely that you would just be fired on the spot, r…
Leave-without-pay normally requires some specific justification(s)/discussion. I've certainly given my manager advanced notice about any longer stretches of vacation and I've tried to do it with awareness of workloads (though for something planned months in advance that's not always possible) but I've pretty much never considered it as asking for permission or it being a negotiation. This is in the US. ADDED: You're…
Where I live that is subject to cancellation of the work contact, you can't lie about why you are absent though the imprisonment can't be cause by itself for laying off.
Re: Re-Evaluating GPT-4's Bar Exam Performance
#97Scoring 96 percentile among humans taking the exam without moving goal posts would have been science fiction two years ago. Now it’s suddenly not good enough and the fact a computer program can score decent among passing lawyers and first time test takers is something to sneer at. The fact I can talk to the computer and it responds to me idiomatically and understands my semantic intent well enough to be nearly indist…
The nerds aren't jaded, they are worried. I'd be too if my job needed nothing more than a keyboard to be completed. There are a lot of people here who need to squeeze another 20-40 years out of a keyboard job.
Re: Re-Evaluating GPT-4's Bar Exam Performance
#98Earlier quoted context omitted.
On any topic that I understand well, LLM output is garbage: it requires more energy to fix it than to solve the original problem to begin with. Are we sure these exams are not present in the training data? (ability to recall information is not impressive for a computer) Still I'm terrible at many many tasks e.g., drawing from description and the models widen significantly types of problems that I can even try (where…
> On any topic that I understand well, LLM output is garbage: it requires more energy to fix it than to solve the original problem to begin with. That's probably true, which is why human most knowledge workers aren't going away any time soon. That said, I have better luck with a different approach: I use LLM's to learn things that I don't already understand well. This forces me to actively understand and validate the…
Re: Re-Evaluating GPT-4's Bar Exam Performance
#99Scoring 96 percentile among humans taking the exam without moving goal posts would have been science fiction two years ago. Now it’s suddenly not good enough and the fact a computer program can score decent among passing lawyers and first time test takers is something to sneer at. The fact I can talk to the computer and it responds to me idiomatically and understands my semantic intent well enough to be nearly indist…
On any topic that I understand well, LLM output is garbage: it requires more energy to fix it than to solve the original problem to begin with. Are we sure these exams are not present in the training data? (ability to recall information is not impressive for a computer) Still I'm terrible at many many tasks e.g., drawing from description and the models widen significantly types of problems that I can even try (where…
Is it generally because the LLM was not trained on that data, therefore have no knowledge of it or because it can't reason well enough?
Re: Re-Evaluating GPT-4's Bar Exam Performance
#100Earlier quoted context omitted.
> which I’d also argue LLMs suck at OK, I’ll bite. What’s your evidence for this argument?
Every bit of interaction I’ve ever had with an LLM. And all the research I’ve seen. They’re plausible word sequence generators, not ‘planning for the future’ agents. Or market analyzers. Or character evaluators. Or anything else. And they tend to be really ‘gullible’. What evidence do you have they could do any of those things? (And not just generate plausible text at a prompt, but actually do those things)
But there's a scarier further step: When people assume an exceptional text-specialist model can also meta-impersonate a generalist model impersonating a specific and different kind of specialist! ("LLM, create a legal defense.")