Live data from Hacker News

Google's Bard shows big leap on LLM performance leaderboard

twitter.com

81–90 of 96 posts

Re: Google's Bard shows big leap on LLM performance leaderboard

#81

Earlier quoted context omitted.

Google is fighting for its life here. They are not worried about cost.

Then what's your explanation for them not deploying Gemini? It's either to expensive or they just don't think it's that important to have a competitive, accessible model out there (which is a valid stance imo).

I don’t really know tbh. I don’t think their current position is because of penny pinching. They invented the transformer and slept on it while OAI ate their lunch.

Re: Google's Bard shows big leap on LLM performance leaderboard

#82

Earlier quoted context omitted.

Bard is free, while GPT-4 is not, so it doesn’t seem like a totally fair comparison. Also what a wild comparison, because afaik chat gpt can’t make charts either.

It absolutely can. Try asking ChatGPT4 something like "Please draw a chart of the population of Switzerland over the past 20 years". It will write Python code to use matplotlib and the resulting chart will appear in the chat. I'd show an example but their sharing feature doesn't work with images for some reason.

> Try asking ChatGPT4 something like "Please draw a chart of the population of Switzerland over the past 20 years".

I tried this on the free version you can get to without paying anything, I didn't realize the advanced version supported this.

Re: Google's Bard shows big leap on LLM performance leaderboard

#83
post #25
post #8

I'm curious about how the benchmark is done. I suspect it can be improved in order to represent user's / usability expectations. I gave Bard a go, after seeing Jeff Dean's tweet. It's just as frustrating as it was, compared to GPT-4. It's simply off the question and unable to realize it's off. I asked it to generate a chart and 3 times it came back with "here's a chart" with no chart, finally saying it doesn't have t…

My favourite by miles was asking bard for some more mathematical examples after explaining quantum field theory quite well in words, to which it said "Ok here are some examples of mathematics: 2+2=4"

https://bard.google.com/share/ff90b8239ebe found it

Re: Google's Bard shows big leap on LLM performance leaderboard

#84

Earlier quoted context omitted.

Lazy??? What does it mean for a model to be lazy?

Q: Please write me a function that makes a couple of text edits and a button. A: Certainly! Here is a function that makes a couple of text edits and a button. function renderForm() { var edit = document.createElement('input'); // The rest of the code is similar. }

I don't get what's lazy about this. Is it because it's not appending the elements?

Re: Google's Bard shows big leap on LLM performance leaderboard

#85
post #38
post #8

I'm curious about how the benchmark is done. I suspect it can be improved in order to represent user's / usability expectations. I gave Bard a go, after seeing Jeff Dean's tweet. It's just as frustrating as it was, compared to GPT-4. It's simply off the question and unable to realize it's off. I asked it to generate a chart and 3 times it came back with "here's a chart" with no chart, finally saying it doesn't have t…

You’re probably using the old Bard model. You can try the new one, bard-jan-24-gemini-pro, by clicking the Direct Chat tab on https://chat.lmsys.org .

Thank you! Quick feedback - lmsys will introduce bias because it lacks rendering support for things like [ x_2 = \frac{1}{2} \left( 1.5 + \frac{2}{1.5} \right) = \frac{1}{2} \left( 1.5 + \frac{4}{3} \right) = \frac{1}{2} \left( \frac{9}{6} + \frac{8}{6} \right) = \frac{17}{12} \approx 1.4166666666666667 ]

Re: Google's Bard shows big leap on LLM performance leaderboard

#86
post #51
post #8

I'm curious about how the benchmark is done. I suspect it can be improved in order to represent user's / usability expectations. I gave Bard a go, after seeing Jeff Dean's tweet. It's just as frustrating as it was, compared to GPT-4. It's simply off the question and unable to realize it's off. I asked it to generate a chart and 3 times it came back with "here's a chart" with no chart, finally saying it doesn't have t…

? https://i.imgur.com/b0fIGyS.png

that makes it even worse for saying 2 times "here's your chart" and the third one appologizing for not having the capability, when in fact is has it?

Here you go https://imgur.com/a/MT04viM

Re: Google's Bard shows big leap on LLM performance leaderboard

#87
post #8

I'm curious about how the benchmark is done. I suspect it can be improved in order to represent user's / usability expectations. I gave Bard a go, after seeing Jeff Dean's tweet. It's just as frustrating as it was, compared to GPT-4. It's simply off the question and unable to realize it's off. I asked it to generate a chart and 3 times it came back with "here's a chart" with no chart, finally saying it doesn't have t…

Bard is free, while GPT-4 is not, so it doesn’t seem like a totally fair comparison. Also what a wild comparison, because afaik chat gpt can’t make charts either.

Are we talking about a capability or pricing benchmark?

Re: Google's Bard shows big leap on LLM performance leaderboard

#88
post #86
post #51

Earlier quoted context omitted.

? https://i.imgur.com/b0fIGyS.png

that makes it even worse for saying 2 times "here's your chart" and the third one appologizing for not having the capability, when in fact is has it? Here you go https://imgur.com/a/MT04viM

That is indeed weird.

Re: Google's Bard shows big leap on LLM performance leaderboard

#89
post #5

how is GPT4-Turbo higher than GPT-4?

It's actually better. All the people claiming unannounced updates make OpenAI model significantly worse were misled by their own lying eyes. It's so hard to believe, I know, but it's true.

ChatGPT was best when first released and absent of censorship ... Probably 10x worst right now

Re: Google's Bard shows big leap on LLM performance leaderboard

#90

for me Bard is much better than ChatGPT... I wish I would be using an uncensored Mistral though.

Yeah, I was evaluating Bard Pro vs Starling and up to the end I wasn't sure what to pick. Starling 7B was almost as good. People don't realize there are ways to

"mommy, can we get GPT?" "we have GPT at home darling"

Post reply on HN