Gemini was a bit disappointing on launch. Good to see them improving. My own experience doesn't really match the arena's leaderboard. I find myself using Claude 2 much more than GPT 4/Turbo. I prefer its out-of-the-box response style, and its answers for my queries seem just as good, if not better. Interesting, Kagi (where I use the chatbots) ranks Claude 1 as equal to (and faster than) GPT 4 (non-Turbo), and marks C…
Google's Bard shows big leap on LLM performance leaderboard
11–20 of 96 posts
Re: Google's Bard shows big leap on LLM performance leaderboard
#12how is GPT4-Turbo higher than GPT-4?
Re: Google's Bard shows big leap on LLM performance leaderboard
#13Getting close, but this will force OpenAI to come out with gpt5 for another big round of catch up
Re: Google's Bard shows big leap on LLM performance leaderboard
#14I'm curious about how the benchmark is done. I suspect it can be improved in order to represent user's / usability expectations. I gave Bard a go, after seeing Jeff Dean's tweet. It's just as frustrating as it was, compared to GPT-4. It's simply off the question and unable to realize it's off. I asked it to generate a chart and 3 times it came back with "here's a chart" with no chart, finally saying it doesn't have t…
Re: Google's Bard shows big leap on LLM performance leaderboard
#15Comparatively BARD models have relatively lower number of votes. I would wait until number of votes are in the same magnitude as other models.
Re: Google's Bard shows big leap on LLM performance leaderboard
#16Comparatively BARD models have relatively lower number of votes. I would wait until number of votes are in the same magnitude as other models.
[0] https://colab.research.google.com/drive/1KdwokPjirkTmpO_P1WB...
Re: Google's Bard shows big leap on LLM performance leaderboard
#17I'm curious about how the benchmark is done. I suspect it can be improved in order to represent user's / usability expectations. I gave Bard a go, after seeing Jeff Dean's tweet. It's just as frustrating as it was, compared to GPT-4. It's simply off the question and unable to realize it's off. I asked it to generate a chart and 3 times it came back with "here's a chart" with no chart, finally saying it doesn't have t…
It’s based on users blind rating the output of two LLMs given the same prompt.
Also I felt I keep getting more personalized results, i.e. models are somehow biased towards user. I heard they plan it, I don't know it's launched, but I feel it.
And there's also fine-tuning in the other direction - my brain got used to ways of interacting with GPT. Same as with Google, I just somehow subconsciously know how to write prompts that get me what I want.
Re: Google's Bard shows big leap on LLM performance leaderboard
#18I'm curious about how the benchmark is done. I suspect it can be improved in order to represent user's / usability expectations. I gave Bard a go, after seeing Jeff Dean's tweet. It's just as frustrating as it was, compared to GPT-4. It's simply off the question and unable to realize it's off. I asked it to generate a chart and 3 times it came back with "here's a chart" with no chart, finally saying it doesn't have t…
If this aint anecdotal evidence, I dont know what is. You need to assume a whole lot to draw these conclusions from prompt output comparison.
Heck, just compare a couple outcomes from ChatGPT and you should conclude the same thing, assuming rationality of course.
Re: Google's Bard shows big leap on LLM performance leaderboard
#19Earlier quoted context omitted.
It’s based on users blind rating the output of two LLMs given the same prompt.
Where GPT really improved for me subjectively is long context - at least since launch. Such ranking can't compare it. Also I felt I keep getting more personalized results, i.e. models are somehow biased towards user. I heard they plan it, I don't know it's launched, but I feel it. And there's also fine-tuning in the other direction - my brain got used to ways of interacting with GPT. Same as with Google, I just someh…