Earlier quoted context omitted.
On any topic that I understand well, LLM output is garbage: it requires more energy to fix it than to solve the original problem to begin with. Are we sure these exams are not present in the training data? (ability to recall information is not impressive for a computer) Still I'm terrible at many many tasks e.g., drawing from description and the models widen significantly types of problems that I can even try (where…
I suspect by garbage you mean not perfect. To be more precise can you please give a topic you know well and your % guess how often the answers are wrong on the topic?
Re-Evaluating GPT-4's Bar Exam Performance
81–90 of 139 posts
Re: Re-Evaluating GPT-4's Bar Exam Performance
#82Earlier quoted context omitted.
the models that you have tried .. are garbage. hmmm Maybe you are not among the many, many, many inside professionals and unofrmed services that have different access than you? money talks?
It is remarkable that folks who tried a garbage LLM like copilot, 3.5, Gemini, or made meta LLMs say naughty words, seem to think these are still SOA. Sometimes I stumble on them and I am shocked at the degradation in quality then realize my settings are wrong. People are vastly underestimating the rate of change here.
It is like a calculator that only worked in one digit, and now it works on 2, the improvement is immense but its still nowhere close to replacing mathematicians since it isn't even working on the same kind of problems.
Edit: In several years we might have a perfect calculator that is better than any human at such tasks, but it still doesn't beat humans at stuff unrelated to calculations. Or in the case of LLMs pattern matching texts, humans don't pattern match texts to plan or mentally simulate scenarios etc, that part isn't covered by LLMs. Human level planning with todays LLM level pattern matching on text would be really useful, we see a lot of humans work that way by using the LLM as a pattern matcher, but there is no progress on automating human level planning so far, LLMs aren't it.
Re: Re-Evaluating GPT-4's Bar Exam Performance
#83Earlier quoted context omitted.
the models that you have tried .. are garbage. hmmm Maybe you are not among the many, many, many inside professionals and unofrmed services that have different access than you? money talks?
It is remarkable that folks who tried a garbage LLM like copilot, 3.5, Gemini, or made meta LLMs say naughty words, seem to think these are still SOA. Sometimes I stumble on them and I am shocked at the degradation in quality then realize my settings are wrong. People are vastly underestimating the rate of change here.
GPT-3.5 was released in March 2022. We are now in June 2024. Over 2 years later.
And on average GPT-4 is about 40% more accurate.
For me, LLMs are very much like self-driving cars. On the journey towards perfect accuracy it gets progressively harder to make advancements.
And for it to replace the status quo it really does need to be perfect. And there is no evidence or research that this is possible.
Re: Re-Evaluating GPT-4's Bar Exam Performance
#84Earlier quoted context omitted.
> The fact I can talk to the computer and it responds to me idiomatically and understands my semantic intent well enough to be nearly indistinguishable from a human being is breath taking That's called a programming language. It's nothing new.
It's a programming language except the programming part, and the language part.
Edit: LLMs biggest feat is being a natural language interpreter, so it can run natural language scripts. It is far from perfect at it, but that is still programming.
Re: Re-Evaluating GPT-4's Bar Exam Performance
#85Scoring 96 percentile among humans taking the exam without moving goal posts would have been science fiction two years ago. Now it’s suddenly not good enough and the fact a computer program can score decent among passing lawyers and first time test takers is something to sneer at. The fact I can talk to the computer and it responds to me idiomatically and understands my semantic intent well enough to be nearly indist…
Re: Re-Evaluating GPT-4's Bar Exam Performance
#86Earlier quoted context omitted.
>You can just ask it, you know. But my question will not be part of the context of that conversation.
Mine was. I asked it the first question, first.
Re: Re-Evaluating GPT-4's Bar Exam Performance
#87Scoring 96 percentile among humans taking the exam without moving goal posts would have been science fiction two years ago. Now it’s suddenly not good enough and the fact a computer program can score decent among passing lawyers and first time test takers is something to sneer at. The fact I can talk to the computer and it responds to me idiomatically and understands my semantic intent well enough to be nearly indist…
But they aren't meaningful for anything other than humans since the correlations between abilities which make them reasonable proxies are not the same.
The idea that these kind of test results prove anything (other than the utility of the tested LLM for humans cheating on the exam) is only valid if you assume not only that the LLM is actually an AGI, but that it's an AGI that is indistinguishable, psychometrically, from a human.
(Which makes a nice circular argument, since these test results are often cited to prove that the LLMs are, or are approaching, AGI.)
Re: Re-Evaluating GPT-4's Bar Exam Performance
#88Earlier quoted context omitted.
On any topic that I understand well, LLM output is garbage: it requires more energy to fix it than to solve the original problem to begin with. Are we sure these exams are not present in the training data? (ability to recall information is not impressive for a computer) Still I'm terrible at many many tasks e.g., drawing from description and the models widen significantly types of problems that I can even try (where…
I suspect by garbage you mean not perfect. To be more precise can you please give a topic you know well and your % guess how often the answers are wrong on the topic?
Re: Re-Evaluating GPT-4's Bar Exam Performance
#89Earlier quoted context omitted.
> On any topic that I understand well, LLM output is garbage: it requires more energy to fix it than to solve the original problem to begin with. That's probably true, which is why human most knowledge workers aren't going away any time soon. That said, I have better luck with a different approach: I use LLM's to learn things that I don't already understand well. This forces me to actively understand and validate the…
I would classify all of those as "non-traditional" learning techniques, unless you actually mean using a textbook while taking a class with a human teacher. Well written textbooks are consumable on their own for some people, but most are not written for that.
Re: Re-Evaluating GPT-4's Bar Exam Performance
#90Scoring 96 percentile among humans taking the exam without moving goal posts would have been science fiction two years ago. Now it’s suddenly not good enough and the fact a computer program can score decent among passing lawyers and first time test takers is something to sneer at. The fact I can talk to the computer and it responds to me idiomatically and understands my semantic intent well enough to be nearly indist…
The real problem is that tests used for humans are callibrated based on the way different human abilities correlate: they aren't objectives themselves, they are convenient proxies. But they aren't meaningful for anything other than humans since the correlations between abilities which make them reasonable proxies are not the same. The idea that these kind of test results prove anything (other than the utility of the…
I've noticed one thing that LLMs seem to have trouble with is going "off task".
There are often very structured evaluation scenarios, with a structured set of items and possible responses (even if defined in a an abstract sense). Performance in those settings is often ok to excellent, but when the test scenario changes, the LLM seems to not be able to recognize it, or fails miserably.
The Obama pictures were a good example of that. Humans could recognize what was going on when the task frame changed, but the AI started to fail miserably.
Me and my friends, similarly, often trick LLMs in interactive tasks by starting to go "off script", where the "script" is some assumption that we're acting in good faith with regard to the task. My guess is humans would have a "WTF?" response, or start to recognize what was happening, but a LLM does not.
In the human realm there's an extra-test world, like you're saying, but for the LLM there's always a test world, and nothing more.
If I'm being honest with myself my guess is a lot of these gaps will be filled over the next decade or so, but there will always be some model boundaries, defined not by the data using to estimate the model, but by the framework the model exists within.