I'm curious about how the benchmark is done. I suspect it can be improved in order to represent user's / usability expectations. I gave Bard a go, after seeing Jeff Dean's tweet. It's just as frustrating as it was, compared to GPT-4. It's simply off the question and unable to realize it's off. I asked it to generate a chart and 3 times it came back with "here's a chart" with no chart, finally saying it doesn't have t…
Google's Bard shows big leap on LLM performance leaderboard
51–60 of 96 posts
Re: Google's Bard shows big leap on LLM performance leaderboard
#52From all free LLMs I find bard to be most useful. Chatgpt 3.5 is not even close and it lazy.
Lazy??? What does it mean for a model to be lazy?
Re: Google's Bard shows big leap on LLM performance leaderboard
#53Re: Google's Bard shows big leap on LLM performance leaderboard
#54Earlier quoted context omitted.
The trick is to access the "bard-jan-24-gemini-pro" model, available in direct chat mode here: https://chat.lmsys.org/ . Significantly better than the prior model.
how odd! What exactly is lmsys using? Some hidden API that google give them so they can have a better ranking there?
I don't know about that second part - but it would make sense that google (and others) may want to use lmsys's arena to benchmark their models.
After all, Human A/B tests are far better then the current automated benchmarks.
I would like more info from lmsys as to how they're accessing these though.
Re: Google's Bard shows big leap on LLM performance leaderboard
#55Earlier quoted context omitted.
The trick is to access the "bard-jan-24-gemini-pro" model, available in direct chat mode here: https://chat.lmsys.org/ . Significantly better than the prior model.
how odd! What exactly is lmsys using? Some hidden API that google give them so they can have a better ranking there?
Re: Google's Bard shows big leap on LLM performance leaderboard
#56Wow. I've suspected for a while that Bard's performance has been limited mostly by cost. Google isn't charging for Bard and they didn't want to run a gigantic model for everyone for free forever. Maybe they made a breakthrough in inference cost for their better models? Or maybe they got tired of everyone clowning on them for being behind and decided to eat the cost for a while. I still think they ought to launch a su…
Everyone else needs to pay nvidia margins.
Training is murkier as it’s more about the total performance and scalability of the system.
Re: Google's Bard shows big leap on LLM performance leaderboard
#57Earlier quoted context omitted.
> Bard is free, while GPT-4 is not, so it doesn’t seem like a totally fair comparison. Wait so if Google suddenly started charging for Bard, it would be instantly better?
It could be. Bard is clearly not the best model Google has and the only reason not to serve the better models is inference cost.
Re: Google's Bard shows big leap on LLM performance leaderboard
#58how is GPT4-Turbo higher than GPT-4?
It's actually better. All the people claiming unannounced updates make OpenAI model significantly worse were misled by their own lying eyes. It's so hard to believe, I know, but it's true.
Re: Google's Bard shows big leap on LLM performance leaderboard
#59Earlier quoted context omitted.
It could be. Bard is clearly not the best model Google has and the only reason not to serve the better models is inference cost.
Google is fighting for its life here. They are not worried about cost.
Re: Google's Bard shows big leap on LLM performance leaderboard
#60From all free LLMs I find bard to be most useful. Chatgpt 3.5 is not even close and it lazy.
Lazy??? What does it mean for a model to be lazy?
For example it might tell me to read some documentation and find the answer myself. So I'm thinking "Yes well, what do I need you for then?"