Earlier quoted context omitted.
It's like watching a baby learn how to talk..
...and saying it would never replace you in your job because he talks like a baby
Asking 60 LLMs a set of 20 questions
181–190 of 352 posts
Re: Asking 60 LLMs a set of 20 questions
#182Re: Asking 60 LLMs a set of 20 questions
#183> Sally (a girl) has 3 brothers. Each brother has 2 sisters. How many sisters does Sally have? The site reports every LLM as getting this wrong. But GPT4 seems to get it right for me: > Sally has 3 brothers. Since each brother has 2 sisters and Sally is one of those sisters, the other sister is the second sister for each brother. So, Sally has 1 sister.
It appears the GPT4 learned it and now it's repeating the correct answer?
Re: Asking 60 LLMs a set of 20 questions
#184Re: Asking 60 LLMs a set of 20 questions
#185What is the point of all these different models? Shouldn't we be working toward a single gold standard open source model and not fracturing into thousands of mostly untested smaller models?
Re: Asking 60 LLMs a set of 20 questions
#186Earlier quoted context omitted.
It might be trained on this question or a variant of it.
It's certainly RLHFed. All of the logic puzzles I use for evaluation that used to fail months ago now pass no problem and I've even had a hard time modifying them to fail.
Q: Bobby (a boy) has 3 sisters. Each sister has 2 brothers. How many brothers does Bobby have? Let's think step by step.
A: First, we know that Bobby has 3 sisters.
Second, we know that each sister has 2 brothers.
This means that Bobby has 2 brothers because the sisters' brothers are Bobby and his two brothers.
So, Bobby has 2 brothers.Re: Asking 60 LLMs a set of 20 questions
#187I have seen numerous posts of llm q&a and by the time people try to replicate them gpt4 is fixed. It either means that OpenAI is actively monitoring the Internet and fixes them or the Internet is actively conspiring to present falsified results for gpt4 to discredit OpenAI
GPT-4 (at least) is explicit in saying that it's learning from user's assessments of its answers, so yes, the only valid way to test is to give it a variation of the prompt and see how well that does. GPT-4 failed the "Sally" test for the first time after 8 tries when I changed every parameter. It got it right on the next try.
Re: Asking 60 LLMs a set of 20 questions
#188Ok, so can we use LLMs to evaluate which LLM performs best on these questions?
Re: Asking 60 LLMs a set of 20 questions
#189Earlier quoted context omitted.
It's certainly RLHFed. All of the logic puzzles I use for evaluation that used to fail months ago now pass no problem and I've even had a hard time modifying them to fail.
And it's only fixed for the stated case, but if you reverse the genders, GPT-4 gets it wrong. Q: Bobby (a boy) has 3 sisters. Each sister has 2 brothers. How many brothers does Bobby have? Let's think step by step. A: First, we know that Bobby has 3 sisters. Second, we know that each sister has 2 brothers. This means that Bobby has 2 brothers because the sisters' brothers are Bobby and his two brothers. So, Bobby has…
Alright, let's break it down step by step:
1. We know Bobby has 3 sisters.
2. Each of those sisters has 2 brothers.
Now, when we say that each sister has 2 brothers, Bobby is one of those brothers. This implies there's another boy in addition to Bobby.
3. That other boy would be Bobby's only brother.
Therefore, Bobby has 1 brother.
Re: Asking 60 LLMs a set of 20 questions
#190is anyone else feeling completely depressed and demotivated by how quickly this is happening?
Not depressed but just makes me question the value of our existence. If we are truly able to create machines with our intellectual potential, where does that leave us? Do we compete with the machines in a rat-race for a chance at happy existence? Or will we create a utopia where the machines do the difficult work and allow for an enjoyable existence for all. The people who are rooting for LLM's are hopefully working…
And you seriously ask, what is there for humans to do? Isn't this the most noble goal there is?