Live data from Hacker News

Google's Bard shows big leap on LLM performance leaderboard

twitter.com

11–20 of 96 posts

Re: Google's Bard shows big leap on LLM performance leaderboard

#11

Gemini was a bit disappointing on launch. Good to see them improving. My own experience doesn't really match the arena's leaderboard. I find myself using Claude 2 much more than GPT 4/Turbo. I prefer its out-of-the-box response style, and its answers for my queries seem just as good, if not better. Interesting, Kagi (where I use the chatbots) ranks Claude 1 as equal to (and faster than) GPT 4 (non-Turbo), and marks C…

Completely concur with your experience. I've always suspected that the questions asked on these arena sites tend towards being a bit different qualitatively than real-world use cases.

Re: Google's Bard shows big leap on LLM performance leaderboard

#13
It’s pretty good got newer info than gpt4, but refuses on stuff that doesn’t make sense when I ask it to write code using a certain public library, it refuses due to privacy reasons. Where as gpt4 will do it.

Getting close, but this will force OpenAI to come out with gpt5 for another big round of catch up

Re: Google's Bard shows big leap on LLM performance leaderboard

#14
post #8

I'm curious about how the benchmark is done. I suspect it can be improved in order to represent user's / usability expectations. I gave Bard a go, after seeing Jeff Dean's tweet. It's just as frustrating as it was, compared to GPT-4. It's simply off the question and unable to realize it's off. I asked it to generate a chart and 3 times it came back with "here's a chart" with no chart, finally saying it doesn't have t…

It’s based on users blind rating the output of two LLMs given the same prompt.

Re: Google's Bard shows big leap on LLM performance leaderboard

#16
post #4

Comparatively BARD models have relatively lower number of votes. I would wait until number of votes are in the same magnitude as other models.

3000 matches should be more than enough for an Elo-style ranking system to converge. But the notebook[0] linked to from the site includes a section with an equal number of matches being sampled for each pairwise comparison, and the results are pretty much the same.

[0] https://colab.research.google.com/drive/1KdwokPjirkTmpO_P1WB...

Re: Google's Bard shows big leap on LLM performance leaderboard

#17
post #8

I'm curious about how the benchmark is done. I suspect it can be improved in order to represent user's / usability expectations. I gave Bard a go, after seeing Jeff Dean's tweet. It's just as frustrating as it was, compared to GPT-4. It's simply off the question and unable to realize it's off. I asked it to generate a chart and 3 times it came back with "here's a chart" with no chart, finally saying it doesn't have t…

It’s based on users blind rating the output of two LLMs given the same prompt.

Where GPT really improved for me subjectively is long context - at least since launch. Such ranking can't compare it.

Also I felt I keep getting more personalized results, i.e. models are somehow biased towards user. I heard they plan it, I don't know it's launched, but I feel it.

And there's also fine-tuning in the other direction - my brain got used to ways of interacting with GPT. Same as with Google, I just somehow subconsciously know how to write prompts that get me what I want.

Re: Google's Bard shows big leap on LLM performance leaderboard

#18
post #8

I'm curious about how the benchmark is done. I suspect it can be improved in order to represent user's / usability expectations. I gave Bard a go, after seeing Jeff Dean's tweet. It's just as frustrating as it was, compared to GPT-4. It's simply off the question and unable to realize it's off. I asked it to generate a chart and 3 times it came back with "here's a chart" with no chart, finally saying it doesn't have t…

> It's just as frustrating as it was, compared to GPT-4. It's simply off the question and unable to realize it's off.

If this aint anecdotal evidence, I dont know what is. You need to assume a whole lot to draw these conclusions from prompt output comparison.

Heck, just compare a couple outcomes from ChatGPT and you should conclude the same thing, assuming rationality of course.

Re: Google's Bard shows big leap on LLM performance leaderboard

#19
post #17

Earlier quoted context omitted.

It’s based on users blind rating the output of two LLMs given the same prompt.

Where GPT really improved for me subjectively is long context - at least since launch. Such ranking can't compare it. Also I felt I keep getting more personalized results, i.e. models are somehow biased towards user. I heard they plan it, I don't know it's launched, but I feel it. And there's also fine-tuning in the other direction - my brain got used to ways of interacting with GPT. Same as with Google, I just someh…

Typically benchmarks have limited aspects they are measuring. I can imagine another suite of benchmarks with longer contexts, but in that case, it might be more difficult to do it in a blind comparison form. At the least, it would be quite costly to run such benchmarks.
Post reply on HN