Live data from Hacker News

Google's Bard shows big leap on LLM performance leaderboard

twitter.com

41–50 of 96 posts

Re: Google's Bard shows big leap on LLM performance leaderboard

#41
post #5

how is GPT4-Turbo higher than GPT-4?

No clue, but it wouldn't surprised me if they identified inputs and training that weren't actually helping. Just as human beings aren't necessarily more helpful by having more varied input, the same seems to apply to LLMs. The interesting thing is it still has a 128k context length. This is awesome because GPT became way more useful to me once it reached this level of context.

Just a caveat, this is 128k context length for ingestion, I believe the output is still constrained to just 4k tokens.

Re: Google's Bard shows big leap on LLM performance leaderboard

#43
post #40

It's not a valid comparison. Bard uses Google (ie the Internet) to answer things, while GPT4 doesn't. So bard can answer things like, "what's the weather in SF today" or "who won the basketball game last night".

I'm not sure what you mean. ChatGPT with GPT4 uses Bing to search the web.

Re: Google's Bard shows big leap on LLM performance leaderboard

#44
post #40

It's not a valid comparison. Bard uses Google (ie the Internet) to answer things, while GPT4 doesn't. So bard can answer things like, "what's the weather in SF today" or "who won the basketball game last night".

Chat GPT4 does do internet searches now.

Re: Google's Bard shows big leap on LLM performance leaderboard

#45
https://twitter.com/JeffDean/status/1750930658900517157

> Bard, powered by the Gemini Pro-scale model, debuts at the #2 position on the independent lmsys leaderboard.

According to Jeff Dean's tweet, it looks like they have a new "Gemini Pro-scale model" being rolled out, not sure what it means by "Pro-scale" though. Also not sure if everyone already got it...

Re: Google's Bard shows big leap on LLM performance leaderboard

#46
post #40

It's not a valid comparison. Bard uses Google (ie the Internet) to answer things, while GPT4 doesn't. So bard can answer things like, "what's the weather in SF today" or "who won the basketball game last night".

Chat GPT4 does do internet searches now.

The model being compared on the LLM Arena does not use ChatGPT4, it uses GPT4 API, which does not have access to the internet by default.

Re: Google's Bard shows big leap on LLM performance leaderboard

#47
post #40

It's not a valid comparison. Bard uses Google (ie the Internet) to answer things, while GPT4 doesn't. So bard can answer things like, "what's the weather in SF today" or "who won the basketball game last night".

I'm not sure what you mean. ChatGPT with GPT4 uses Bing to search the web.

See reply to sibling comment.

Re: Google's Bard shows big leap on LLM performance leaderboard

#49
post #34
post #9

Wow. I've suspected for a while that Bard's performance has been limited mostly by cost. Google isn't charging for Bard and they didn't want to run a gigantic model for everyone for free forever. Maybe they made a breakthrough in inference cost for their better models? Or maybe they got tired of everyone clowning on them for being behind and decided to eat the cost for a while. I still think they ought to launch a su…

The trick is to access the "bard-jan-24-gemini-pro" model, available in direct chat mode here: https://chat.lmsys.org/ . Significantly better than the prior model.

how odd! What exactly is lmsys using? Some hidden API that google give them so they can have a better ranking there?
Post reply on HN