Live data from Hacker News

Google's Bard shows big leap on LLM performance leaderboard

twitter.com

1–10 of 96 posts

Re: Google's Bard shows big leap on LLM performance leaderboard

#6

this leaderboard seems easily cheated/gamed. once enough eyes are on it it will be worthless

Kind of, ChatBot Arena was updated on Dec 7th, so hard to know for sure if Bard is just fine tuned against it.

That being said Chatbot Arena is a pretty wide variety of test scenarios. If fine tuning made the model perfect, all the small models would get similar scores to GPT4, which they don't. Essentially it ranks how people believe a ChatBot should respond, rather than just zero shot, 1 shot and COT type benchmarks.

Re: Google's Bard shows big leap on LLM performance leaderboard

#7
Gemini was a bit disappointing on launch. Good to see them improving.

My own experience doesn't really match the arena's leaderboard. I find myself using Claude 2 much more than GPT 4/Turbo. I prefer its out-of-the-box response style, and its answers for my queries seem just as good, if not better.

Interesting, Kagi (where I use the chatbots) ranks Claude 1 as equal to (and faster than) GPT 4 (non-Turbo), and marks Claude 2's quality as on-par with 4 Turbo (albeit slower). Worth noting it's a simple 1-4-star ranking, unlike the arena's ELO numbers.

Re: Google's Bard shows big leap on LLM performance leaderboard

#8
I'm curious about how the benchmark is done. I suspect it can be improved in order to represent user's / usability expectations.

I gave Bard a go, after seeing Jeff Dean's tweet.

It's just as frustrating as it was, compared to GPT-4. It's simply off the question and unable to realize it's off.

I asked it to generate a chart and 3 times it came back with "here's a chart" with no chart, finally saying it doesn't have that ability.

Re: Google's Bard shows big leap on LLM performance leaderboard

#9
Wow. I've suspected for a while that Bard's performance has been limited mostly by cost. Google isn't charging for Bard and they didn't want to run a gigantic model for everyone for free forever. Maybe they made a breakthrough in inference cost for their better models? Or maybe they got tired of everyone clowning on them for being behind and decided to eat the cost for a while.

I still think they ought to launch a subscription so we can see their absolute best model running in public.

Post reply on HN