Live data from Hacker News

Google's Bard shows big leap on LLM performance leaderboard

twitter.com

31–40 of 96 posts

Re: Google's Bard shows big leap on LLM performance leaderboard

#31

Anyone know the difference between Bard Gemini Pro and Gemini Pro Dev API on the leaderboard?

While the other response was unflattering but true, a better answer would be that Bard ___ is trained to be a general purpose chat bot, specific for the Bard user experience. While Gemini ___ API is a developer focused API for LLM use. Similar to how ChatGPT performs different to the OpenAI APIs or Copilot in Bing, or whatever. Basically, they’re fine tuned and prompted to respond differently.

Re: Google's Bard shows big leap on LLM performance leaderboard

#33
post #8

I'm curious about how the benchmark is done. I suspect it can be improved in order to represent user's / usability expectations. I gave Bard a go, after seeing Jeff Dean's tweet. It's just as frustrating as it was, compared to GPT-4. It's simply off the question and unable to realize it's off. I asked it to generate a chart and 3 times it came back with "here's a chart" with no chart, finally saying it doesn't have t…

Bard is free, while GPT-4 is not, so it doesn’t seem like a totally fair comparison. Also what a wild comparison, because afaik chat gpt can’t make charts either.

> Bard is free, while GPT-4 is not, so it doesn’t seem like a totally fair comparison.

Wait so if Google suddenly started charging for Bard, it would be instantly better?

Re: Google's Bard shows big leap on LLM performance leaderboard

#34
post #9

Wow. I've suspected for a while that Bard's performance has been limited mostly by cost. Google isn't charging for Bard and they didn't want to run a gigantic model for everyone for free forever. Maybe they made a breakthrough in inference cost for their better models? Or maybe they got tired of everyone clowning on them for being behind and decided to eat the cost for a while. I still think they ought to launch a su…

The trick is to access the "bard-jan-24-gemini-pro" model, available in direct chat mode here: https://chat.lmsys.org/. Significantly better than the prior model.

Re: Google's Bard shows big leap on LLM performance leaderboard

#37
post #33

Earlier quoted context omitted.

Bard is free, while GPT-4 is not, so it doesn’t seem like a totally fair comparison. Also what a wild comparison, because afaik chat gpt can’t make charts either.

> Bard is free, while GPT-4 is not, so it doesn’t seem like a totally fair comparison. Wait so if Google suddenly started charging for Bard, it would be instantly better?

It could be. Bard is clearly not the best model Google has and the only reason not to serve the better models is inference cost.

Re: Google's Bard shows big leap on LLM performance leaderboard

#38
post #8

I'm curious about how the benchmark is done. I suspect it can be improved in order to represent user's / usability expectations. I gave Bard a go, after seeing Jeff Dean's tweet. It's just as frustrating as it was, compared to GPT-4. It's simply off the question and unable to realize it's off. I asked it to generate a chart and 3 times it came back with "here's a chart" with no chart, finally saying it doesn't have t…

You’re probably using the old Bard model. You can try the new one, bard-jan-24-gemini-pro, by clicking the Direct Chat tab on https://chat.lmsys.org.

Re: Google's Bard shows big leap on LLM performance leaderboard

#39
post #5

how is GPT4-Turbo higher than GPT-4?

No clue, but it wouldn't surprised me if they identified inputs and training that weren't actually helping. Just as human beings aren't necessarily more helpful by having more varied input, the same seems to apply to LLMs.

The interesting thing is it still has a 128k context length. This is awesome because GPT became way more useful to me once it reached this level of context.

Post reply on HN