> Claude Instant v1 > Sally has 0 sisters. The question provides no information about Sally having any sisters herself. It isn't entirely wrong, is it?
Asking 60 LLMs a set of 20 questions
91–100 of 352 posts
Re: Asking 60 LLMs a set of 20 questions
#92Re: Asking 60 LLMs a set of 20 questions
#93Earlier quoted context omitted.
Nondeterminism strikes again! But yes, I would expect GPT-4 to get this right most of the time.
Saying "Sorry, I was non-deterministic" to your teacher won't do much for your grade.
Re: Asking 60 LLMs a set of 20 questions
#94I get frustrated when I tell an LLM “reply only with x” and then rather than responding “x”, it still responds with “Sure thing! Here’s x” or some other extra words.
Re: Asking 60 LLMs a set of 20 questions
#95Has anyone looked through all the responses and chosen any winners?
GPT4 seems to me to be the best. Undi95/ReMM-SLERP-L2-13B the runner up.
Re: Asking 60 LLMs a set of 20 questions
#96Earlier quoted context omitted.
Humans. After all, LLMs are designed to reason equal to or better than humans.
By "Humans", I assume you mean something like "adult humans, well-educated in the relevant fields". Otherwise, most of these responses look like they would easily beat most humans.
Me, Kubernetes Haikus, time taken 84 seconds:
----------
Kubernetes rules
With its smooth orchestration
You can reach web scale
----------
Kubernetes sucks
Lost in endless YAML hell
Why is it broken?
Re: Asking 60 LLMs a set of 20 questions
#97> Sally (a girl) has 3 brothers. Each brother has 2 sisters. How many sisters does Sally have? The site reports every LLM as getting this wrong. But GPT4 seems to get it right for me: > Sally has 3 brothers. Since each brother has 2 sisters and Sally is one of those sisters, the other sister is the second sister for each brother. So, Sally has 1 sister.
Am I wrong to think that? Are LLMs in the future going to be able to “think through” actual logic problems?
Re: Asking 60 LLMs a set of 20 questions
#98Earlier quoted context omitted.
What alternative technology do you think is better? In other words, what is your frame of reference for labeling this "pretty terrible"?
Given that people are already firing real human workers to replace them with worse but cheaper LLMs, I'd argue that we're not talking about a competing technology, but that the competition is simply not firing your workforce. And, as an obligate customer of many large companies, you should be in favor of that as well. Most companies already automate, poorly, a great deal of customer service work; let us hope they do…
Re: Asking 60 LLMs a set of 20 questions
#99> Sally (a girl) has 3 brothers. Each brother has 2 sisters. How many sisters does Sally have? The site reports every LLM as getting this wrong. But GPT4 seems to get it right for me: > Sally has 3 brothers. Since each brother has 2 sisters and Sally is one of those sisters, the other sister is the second sister for each brother. So, Sally has 1 sister.
I wouldn’t expect an LLM to get this right unless it had been trained on a solution. Am I wrong to think that? Are LLMs in the future going to be able to “think through” actual logic problems?
Re: Asking 60 LLMs a set of 20 questions
#100> Sally (a girl) has 3 brothers. Each brother has 2 sisters. How many sisters does Sally have? The site reports every LLM as getting this wrong. But GPT4 seems to get it right for me: > Sally has 3 brothers. Since each brother has 2 sisters and Sally is one of those sisters, the other sister is the second sister for each brother. So, Sally has 1 sister.
I wouldn’t expect an LLM to get this right unless it had been trained on a solution. Am I wrong to think that? Are LLMs in the future going to be able to “think through” actual logic problems?