Live data from Hacker News

An In-depth Look at Gemini's Language Abilities

arxiv.org

21–30 of 73 posts

Re: An In-depth Look at Gemini's Language Abilities

#21

It's incredible how accurate the Chatbot Arena Leaderboard [0] is at predicting model performance compared to benchmarks (which can and are being gamed, see all the 7B models on HF leaderboard) [0]: https://huggingface.co/spaces/lmsys/chatbot-arena-leaderboar...

It's because it isn't "predicting" anything, but rather aggregating user feedback. That is of course going to be closest to judging the subjective "best" model that pleases most people.

It's like saying how can evaluating 5 years of performance at work be better at predicting someone's competency than their SAT scores.

Re: An In-depth Look at Gemini's Language Abilities

#23

The Gemini white paper reports higher scores on HumanEval and other tasks. So one of Google lied, this eval has bugs, they borked the deployment is true

Most likely Google has lied. AI playing video games and board games don't translate to real world applications. Many people fail to see that.

Re: An In-depth Look at Gemini's Language Abilities

#24
post #18
post #13

Earlier quoted context omitted.

It's much more accurate than the Open LLM Leaderboard, that's for sure. Human evaluation has always been the gold standard. I just wish we could filter by the votes which were made after only one or two prompts and I hope they don't include the non-blind votes in the results.

These are the rules of the battle arena: -Ask any question to two anonymous models (e.g., ChatGPT, Claude, Llama) and vote for the better one! -You can continue chatting until you identify a winner. -Vote won’t be counted if model identity is revealed during conversation.

It's not completely blind/anonymous, since you can just ask "What's your name" and the model will identify itself.

Edit: I missed the third rule. I wonder how smart their detection is.

Re: An In-depth Look at Gemini's Language Abilities

#25

I don't understand why people keep falling for Google's ad campaign. Google have its lead in AI playing video games and board games. It is cool, entertaining and all that jazz. But OpenAI and MS are the real leaders in real AI.

https://arxiv.org/abs/1706.03762 https://research.google/pubs/attention-is-all-you-need/

let's not forget where this breakthrough came from, i wouldn't count Google out

Re: An In-depth Look at Gemini's Language Abilities

#26
post #12
post #9

Earlier quoted context omitted.

Mixtral is a mystery to me. How in the world is that team on par with/beating GOOGLE, who presumably have all the resources in the world to throw at this?

Mixtral is on-par with Gemini Pro, not Gemini Ultra (and even there it is further behind Gemini Pro than Gemini Pro is behind GPT 3.5). But to directly answer your question, they are quite well-funded, having raised over $700mil to date. I definitely wouldn't count them out.

Gemini Ultra is not out yet. With the same logic, you could compare an unreleased Mistral model with Gemini Ultra.

Re: An In-depth Look at Gemini's Language Abilities

#27
post #9
post #3

Has anyone (outside of Google) gotten to play with Gemini Ultra yet? Been hearing a lot about Pro, but I'd be interested in seeing whether Ultra is really close to as capable as they claim. Also very interesting that Mixtral 8x7B ranks in the same neighborhood as Gemini Pro/GPT 3.5 Turbo/Claude 2.1 while being fully open source and Apache 2.0 licensed.

Mixtral is a mystery to me. How in the world is that team on par with/beating GOOGLE, who presumably have all the resources in the world to throw at this?

There's a survivorship bias going on here. You've never heard of the thousands of teams out there that are Mistral's size but AREN'T getting results that compete on the global stage, but they do exist. But you've heard of Google, whether they're getting it right or not.

Re: An In-depth Look at Gemini's Language Abilities

#28

Earlier quoted context omitted.

It's astounding that Mixtral Instruct ties with 3.5-turbo while being ~10x smaller.

3.5-turbo might be 20B, not 10x larger https://www.reddit.com/r/LocalLLaMA/comments/17jrj82/new_mic...

Hmm right, the ~300B figure may have been for the non-turbo 3.5

Re: An In-depth Look at Gemini's Language Abilities

#30
post #9
post #3

Has anyone (outside of Google) gotten to play with Gemini Ultra yet? Been hearing a lot about Pro, but I'd be interested in seeing whether Ultra is really close to as capable as they claim. Also very interesting that Mixtral 8x7B ranks in the same neighborhood as Gemini Pro/GPT 3.5 Turbo/Claude 2.1 while being fully open source and Apache 2.0 licensed.

Mixtral is a mystery to me. How in the world is that team on par with/beating GOOGLE, who presumably have all the resources in the world to throw at this?

Mistral.AI was founded by three people from Deepmind, they're beating Google because Google no longer has them.
Post reply on HN