Live data from Hacker News

Google's Bard shows big leap on LLM performance leaderboard

twitter.com

21–30 of 96 posts

Re: Google's Bard shows big leap on LLM performance leaderboard

#23

this leaderboard seems easily cheated/gamed. once enough eyes are on it it will be worthless

I don't think Google would specifically targeting this leaderboard just for marketing purpose though.

Why not? Their entire presentation a month ago was fake.

Re: Google's Bard shows big leap on LLM performance leaderboard

#24
post #8

I'm curious about how the benchmark is done. I suspect it can be improved in order to represent user's / usability expectations. I gave Bard a go, after seeing Jeff Dean's tweet. It's just as frustrating as it was, compared to GPT-4. It's simply off the question and unable to realize it's off. I asked it to generate a chart and 3 times it came back with "here's a chart" with no chart, finally saying it doesn't have t…

Bard is free, while GPT-4 is not, so it doesn’t seem like a totally fair comparison.

Also what a wild comparison, because afaik chat gpt can’t make charts either.

Re: Google's Bard shows big leap on LLM performance leaderboard

#25
post #8

I'm curious about how the benchmark is done. I suspect it can be improved in order to represent user's / usability expectations. I gave Bard a go, after seeing Jeff Dean's tweet. It's just as frustrating as it was, compared to GPT-4. It's simply off the question and unable to realize it's off. I asked it to generate a chart and 3 times it came back with "here's a chart" with no chart, finally saying it doesn't have t…

My favourite by miles was asking bard for some more mathematical examples after explaining quantum field theory quite well in words, to which it said "Ok here are some examples of mathematics: 2+2=4"

Re: Google's Bard shows big leap on LLM performance leaderboard

#26
post #6

this leaderboard seems easily cheated/gamed. once enough eyes are on it it will be worthless

Kind of, ChatBot Arena was updated on Dec 7th, so hard to know for sure if Bard is just fine tuned against it. That being said Chatbot Arena is a pretty wide variety of test scenarios. If fine tuning made the model perfect, all the small models would get similar scores to GPT4, which they don't. Essentially it ranks how people believe a ChatBot should respond, rather than just zero shot, 1 shot and COT type benchmark…

[deleted]

Re: Google's Bard shows big leap on LLM performance leaderboard

#27
post #23

Earlier quoted context omitted.

I don't think Google would specifically targeting this leaderboard just for marketing purpose though.

Why not? Their entire presentation a month ago was fake.

Because this leaderboard has a very limited audience, and second only to GPT-4 turbo isn't good marketing regardless.

Re: Google's Bard shows big leap on LLM performance leaderboard

#28
post #9

Wow. I've suspected for a while that Bard's performance has been limited mostly by cost. Google isn't charging for Bard and they didn't want to run a gigantic model for everyone for free forever. Maybe they made a breakthrough in inference cost for their better models? Or maybe they got tired of everyone clowning on them for being behind and decided to eat the cost for a while. I still think they ought to launch a su…

I think their play is the same as always

it's better to let more people interact with it because this will help training the model (get more data) so it must be free to use.

Re: Google's Bard shows big leap on LLM performance leaderboard

#30

Gemini was a bit disappointing on launch. Good to see them improving. My own experience doesn't really match the arena's leaderboard. I find myself using Claude 2 much more than GPT 4/Turbo. I prefer its out-of-the-box response style, and its answers for my queries seem just as good, if not better. Interesting, Kagi (where I use the chatbots) ranks Claude 1 as equal to (and faster than) GPT 4 (non-Turbo), and marks C…

I rarely hear people talk about liking Claude. I’m curious, what do you use it for?

I personally use bard for general internet search stuff, because I like how google has set up the linking (and the UI). I use GPT-* for technical stuff (eg GitHub codepilot) because I find it can synthesize code better.

Post reply on HN