Live data from Hacker News

Google's Bard shows big leap on LLM performance leaderboard

twitter.com

51–60 of 96 posts

Re: Google's Bard shows big leap on LLM performance leaderboard

#51
post #8

I'm curious about how the benchmark is done. I suspect it can be improved in order to represent user's / usability expectations. I gave Bard a go, after seeing Jeff Dean's tweet. It's just as frustrating as it was, compared to GPT-4. It's simply off the question and unable to realize it's off. I asked it to generate a chart and 3 times it came back with "here's a chart" with no chart, finally saying it doesn't have t…

?

https://i.imgur.com/b0fIGyS.png

Re: Google's Bard shows big leap on LLM performance leaderboard

#52

From all free LLMs I find bard to be most useful. Chatgpt 3.5 is not even close and it lazy.

Lazy??? What does it mean for a model to be lazy?

Does not completely answer the question, all aspects of a question, or has a response otherwise cut off. This was common for gpt4 and coding questions it would simply stop responding in the middl

Re: Google's Bard shows big leap on LLM performance leaderboard

#54
post #49
post #34

Earlier quoted context omitted.

The trick is to access the "bard-jan-24-gemini-pro" model, available in direct chat mode here: https://chat.lmsys.org/ . Significantly better than the prior model.

how odd! What exactly is lmsys using? Some hidden API that google give them so they can have a better ranking there?

> Some hidden API that google give them so they can have a better ranking there?

I don't know about that second part - but it would make sense that google (and others) may want to use lmsys's arena to benchmark their models.

After all, Human A/B tests are far better then the current automated benchmarks.

I would like more info from lmsys as to how they're accessing these though.

Re: Google's Bard shows big leap on LLM performance leaderboard

#55
post #49
post #34

Earlier quoted context omitted.

The trick is to access the "bard-jan-24-gemini-pro" model, available in direct chat mode here: https://chat.lmsys.org/ . Significantly better than the prior model.

how odd! What exactly is lmsys using? Some hidden API that google give them so they can have a better ranking there?

Most likely through this platform: https://console.cloud.google.com/vertex-ai

Re: Google's Bard shows big leap on LLM performance leaderboard

#56
post #9

Wow. I've suspected for a while that Bard's performance has been limited mostly by cost. Google isn't charging for Bard and they didn't want to run a gigantic model for everyone for free forever. Maybe they made a breakthrough in inference cost for their better models? Or maybe they got tired of everyone clowning on them for being behind and decided to eat the cost for a while. I still think they ought to launch a su…

Google do have an inference advantage with TPUs.

Everyone else needs to pay nvidia margins.

Training is murkier as it’s more about the total performance and scalability of the system.

Re: Google's Bard shows big leap on LLM performance leaderboard

#57
post #33

Earlier quoted context omitted.

> Bard is free, while GPT-4 is not, so it doesn’t seem like a totally fair comparison. Wait so if Google suddenly started charging for Bard, it would be instantly better?

It could be. Bard is clearly not the best model Google has and the only reason not to serve the better models is inference cost.

Google is fighting for its life here. They are not worried about cost.

Re: Google's Bard shows big leap on LLM performance leaderboard

#58
post #5

how is GPT4-Turbo higher than GPT-4?

It's actually better. All the people claiming unannounced updates make OpenAI model significantly worse were misled by their own lying eyes. It's so hard to believe, I know, but it's true.

It isn’t, because performance isn’t scalar. It seems to be a more preferable chatbot in this arena. It is objectively less capable for many other domain specific tasks.

Re: Google's Bard shows big leap on LLM performance leaderboard

#59

Earlier quoted context omitted.

It could be. Bard is clearly not the best model Google has and the only reason not to serve the better models is inference cost.

Google is fighting for its life here. They are not worried about cost.

Google is one of the most cost conscious companies in tech when it comes to compute costs. Sure they are blind to other kinds of cost like reputation damage due to stupid leadership decisions. But in terms of their tech, they run their servers and their network to very high utilization, definitely exceeding competitors like Amazon.

Re: Google's Bard shows big leap on LLM performance leaderboard

#60

From all free LLMs I find bard to be most useful. Chatgpt 3.5 is not even close and it lazy.

Lazy??? What does it mean for a model to be lazy?

Sometimes I'd have GPT tell me to how to do something when it could have done it for me.

For example it might tell me to read some documentation and find the answer myself. So I'm thinking "Yes well, what do I need you for then?"

Post reply on HN