Earlier quoted context omitted.
It's a programming language except the programming part, and the language part.
How is the text you write not a language, and how is writing instructions that computers follow not programming? Edit: LLMs biggest feat is being a natural language interpreter, so it can run natural language scripts. It is far from perfect at it, but that is still programming.
Re-Evaluating GPT-4's Bar Exam Performance
121–130 of 139 posts
Re: Re-Evaluating GPT-4's Bar Exam Performance
#122Earlier quoted context omitted.
On what topics you understand well does GOT-4o or Claude Opus produce garbage?
Anything deeper than surface level in medicine. Try getting it to properly select crystalloids with proper additives for a patient with a given history and lab results and watch in horror as it confidently gives instructions that would kill the patient. What is even more irritating is that I had gpt4 debate me on things that it was completely wrong about and it was only when I responded with a stern rebuke that it hi…
Re: Re-Evaluating GPT-4's Bar Exam Performance
#123Earlier quoted context omitted.
> People are vastly underestimating the rate of change here GPT-3.5 was released in March 2022. We are now in June 2024. Over 2 years later. And on average GPT-4 is about 40% more accurate. For me, LLMs are very much like self-driving cars. On the journey towards perfect accuracy it gets progressively harder to make advancements. And for it to replace the status quo it really does need to be perfect. And there is no…
Its enough to decrease the amount of ppl you need in IT by a factor of 20-30%. Ppl dont want to hear that, but you see less and less offers and not only for junior positions. Hard truth is that like with any tool/automation - the higher performance improves, the less ppl are needed for this kind of work. Just look at how some parts of manual labor were made redundant. Why ppl think it wont be the same with mental wor…
E.g. I had it autocompleting a set of 20 variable#s today Something like output.blah=tostring(input[blah]). The kind of work you give to a regex.
In the middle of the list, it decides to go output.blah=some long weitd piece of code, completely unexpected and syntactically invalid.
I am still in my AI evaluation phase, and sometimes I am impressed with what it does. But just as possible is an unexpected total failure. As long is it does that, I can't trust it.
Re: Re-Evaluating GPT-4's Bar Exam Performance
#124Earlier quoted context omitted.
Obviously you need subject knowledge, that should be implicit? Keep in mind even today[1] ( in California and few other states) you don't need to go law school to write the Bar exam and practice law, various forms of apprenticeship under a judge or lawyer are allowed You also don't need to write the exam to practice many aspects of the legal profession. The exam is never meant to be a high bar of quality or selection…
> Obviously you need subject knowledge, that should be implicit? Well, in a lot of the so-called soft sciences, you can easily beat a test without subject knowledge. I had figured that the bar exam might be something like that -- but it's more akin to something like biology, where there are a lot of arcane and counterintuitive little rules that have emerged over time. And you need to know those , or you're sunk. You…
Reminds me of the mandatory trainings you take for work every year. Normal logical thinking can get you through most of them.
Re: Re-Evaluating GPT-4's Bar Exam Performance
#125Earlier quoted context omitted.
it is called a legal code after all
“Code” in that sense predates pretty much any form of computer or technical writing. It came from the same word in old French in the 14th century, which itself came from the Latin codex . It basically meant “book”. Now of course it is specific to books that contain laws.
That's still cognate with the concept of computer code.
Re: Re-Evaluating GPT-4's Bar Exam Performance
#126> Furthermore, unlike its documentation for the other exams it tested (OpenAI 2023b, p. 25), OpenAI’s technical report provides no direct citation for how the UBE percentile was computed, creating further uncertainty over both the original source and validity of the 90th percentile claim. This is the part that bothered me (licensed attorney) from the start. If it scores this high, where are the receipts? I’m sure Ope…
I'm not a licensed attorney, but that's also bothered me about all of these sorts of stories. There is never any proof provided for any of the claims, and the behavior often contradicts what can be observed using the system yourself. I also assume they cook the books a little by having included a bunch of bar exam specific training when creating the model in first place specifically to better on bar exams than in general.
Re: Re-Evaluating GPT-4's Bar Exam Performance
#127Earlier quoted context omitted.
On any topic that I understand well, LLM output is garbage: it requires more energy to fix it than to solve the original problem to begin with. Are we sure these exams are not present in the training data? (ability to recall information is not impressive for a computer) Still I'm terrible at many many tasks e.g., drawing from description and the models widen significantly types of problems that I can even try (where…
>On any topic that I understand well, LLM output is garbage: it requires more energy to fix it than to solve the original problem to begin with. Is it generally because the LLM was not trained on that data, therefore have no knowledge of it or because it can't reason well enough?
Re: Re-Evaluating GPT-4's Bar Exam Performance
#128Earlier quoted context omitted.
If that chatbot had Steve Jobs voice ? I dunno, I would probably prefer to work under that chatbot than my current CEO that only tries to squize as much as possible out of ppl already working for him.
Like the chatbot wouldn’t squeeze you 10x harder. At least a human CEO has to worry about being arrested or someone setting their house on fire.
Re: Re-Evaluating GPT-4's Bar Exam Performance
#129Earlier quoted context omitted.
Anything deeper than surface level in medicine. Try getting it to properly select crystalloids with proper additives for a patient with a given history and lab results and watch in horror as it confidently gives instructions that would kill the patient. What is even more irritating is that I had gpt4 debate me on things that it was completely wrong about and it was only when I responded with a stern rebuke that it hi…
LLMs are not good at answering expert level questions at the forefront of human knowledge.
Re: Re-Evaluating GPT-4's Bar Exam Performance
#130Earlier quoted context omitted.
Specifics like this make it much easier to agree on LLM capabilities, thank you. Automatic proof generation is a massive open problem in all of computer science and not close to be solved. It’s true LLMs aren’t great at it and more is required for example as with the geometry system Deepmind progresses on. On the other hand they can be very useful to explain concepts and allow interactive questioning to drill down an…
How do yo debug its hallucination misinformation via voice interface while you commute?