Asking 60 LLMs a set of 20 questions
41–50 of 352 posts
Re: Asking 60 LLMs a set of 20 questions
#42> Sally (a girl) has 3 brothers. Each brother has 2 sisters. How many sisters does Sally have? The site reports every LLM as getting this wrong. But GPT4 seems to get it right for me: > Sally has 3 brothers. Since each brother has 2 sisters and Sally is one of those sisters, the other sister is the second sister for each brother. So, Sally has 1 sister.
Re: Asking 60 LLMs a set of 20 questions
#43How did you run the queries against these engines? Did you host the inference engines yourself or did you have to sign up for services. If there was a way to supplement each LLM with additional data I can see this being a useful service for companies who are investigating ML in various facets of their business.
Re: Asking 60 LLMs a set of 20 questions
#44Re: Asking 60 LLMs a set of 20 questions
#45Despite the hype about LLMs, many of the answers are pretty terrible. The 12-bar blues progressions seem mostly clueless. The question is will any of these ever get significantly better with time, or are they mostly going to stagnate?
Re: Asking 60 LLMs a set of 20 questions
#46I feel like this bot mocking us
Re: Asking 60 LLMs a set of 20 questions
#47Are these LLMs deterministic or is this comparison rather useless?
Re: Asking 60 LLMs a set of 20 questions
#48> Sally (a girl) has 3 brothers. Each brother has 2 sisters. How many sisters does Sally have? The site reports every LLM as getting this wrong. But GPT4 seems to get it right for me: > Sally has 3 brothers. Since each brother has 2 sisters and Sally is one of those sisters, the other sister is the second sister for each brother. So, Sally has 1 sister.
With the simpler prompt, all the answers were wrong, most of them ridiculously wrong.
Re: Asking 60 LLMs a set of 20 questions
#49Despite the hype about LLMs, many of the answers are pretty terrible. The 12-bar blues progressions seem mostly clueless. The question is will any of these ever get significantly better with time, or are they mostly going to stagnate?
What alternative technology do you think is better? In other words, what is your frame of reference for labeling this "pretty terrible"?
Re: Asking 60 LLMs a set of 20 questions
#50Has anyone looked through all the responses and chosen any winners?
document.querySelectorAll("td pre").forEach((node) => { let code = node.textContent; node.insertAdjacentHTML('afterend', code) })
Or take a look at my screenshot: https://i.ibb.co/Kw0kp58/Screenshot-2023-09-09-at-17-15-20-h...