Has anyone (outside of Google) gotten to play with Gemini Ultra yet? Been hearing a lot about Pro, but I'd be interested in seeing whether Ultra is really close to as capable as they claim. Also very interesting that Mixtral 8x7B ranks in the same neighborhood as Gemini Pro/GPT 3.5 Turbo/Claude 2.1 while being fully open source and Apache 2.0 licensed.
An In-depth Look at Gemini's Language Abilities
31–40 of 73 posts
Re: An In-depth Look at Gemini's Language Abilities
#32Earlier quoted context omitted.
These are the rules of the battle arena: -Ask any question to two anonymous models (e.g., ChatGPT, Claude, Llama) and vote for the better one! -You can continue chatting until you identify a winner. -Vote won’t be counted if model identity is revealed during conversation.
It's not completely blind/anonymous, since you can just ask "What's your name" and the model will identify itself. Edit: I missed the third rule. I wonder how smart their detection is.
Re: An In-depth Look at Gemini's Language Abilities
#33If I was already using GCP and they reduced their price (>10%) and offered tight integration with rest of GCP services it would still be appealing.
Re: An In-depth Look at Gemini's Language Abilities
#34It's incredible how accurate the Chatbot Arena Leaderboard [0] is at predicting model performance compared to benchmarks (which can and are being gamed, see all the 7B models on HF leaderboard) [0]: https://huggingface.co/spaces/lmsys/chatbot-arena-leaderboar...
I only make technical (pytorch) questions though.
Re: An In-depth Look at Gemini's Language Abilities
#35Earlier quoted context omitted.
Mixtral is on-par with Gemini Pro, not Gemini Ultra (and even there it is further behind Gemini Pro than Gemini Pro is behind GPT 3.5). But to directly answer your question, they are quite well-funded, having raised over $700mil to date. I definitely wouldn't count them out.
Gemini Ultra is not out yet. With the same logic, you could compare an unreleased Mistral model with Gemini Ultra.
Surely you could make a comparison of two unreleased models, but it wouldn't be interesting because you don't have any real data (and benchmarks don't really mean anything).
Re: An In-depth Look at Gemini's Language Abilities
#36I don't understand why people keep falling for Google's ad campaign. Google have its lead in AI playing video games and board games. It is cool, entertaining and all that jazz. But OpenAI and MS are the real leaders in real AI.
Re: An In-depth Look at Gemini's Language Abilities
#37Earlier quoted context omitted.
Mixtral is a mystery to me. How in the world is that team on par with/beating GOOGLE, who presumably have all the resources in the world to throw at this?
There's a survivorship bias going on here. You've never heard of the thousands of teams out there that are Mistral's size but AREN'T getting results that compete on the global stage, but they do exist. But you've heard of Google, whether they're getting it right or not.
Re: An In-depth Look at Gemini's Language Abilities
#38It's incredible how accurate the Chatbot Arena Leaderboard [0] is at predicting model performance compared to benchmarks (which can and are being gamed, see all the 7B models on HF leaderboard) [0]: https://huggingface.co/spaces/lmsys/chatbot-arena-leaderboar...
Thanks for the reference I was searching for a benchmark that can quantify the typical user experience, as most synthetic ones are completly ineffective. At what sample size the ranking become significant? Or is it baked in the metrics (ELO)?
The Glicko rating system is very similar to Elo, but it also models the variance of a given rating. It can directly tell you a "rating deviation."
Re: An In-depth Look at Gemini's Language Abilities
#39I don't understand why people keep falling for Google's ad campaign. Google have its lead in AI playing video games and board games. It is cool, entertaining and all that jazz. But OpenAI and MS are the real leaders in real AI.
Even if you don't think Google doesn't have a talent or product chops to be leaders in AI, Google can do things cheaper than others because of their infrastructure. When they do release something useful they'll probably be able to offer it free and force it on people on the most visited pages/most used browser. Surprised how many people think having a years head start means OpenAI and Microsoft are going to always be…
Why do you way they will 'probably' do that? Do you have any information to back that up or is this your speculation?
Re: An In-depth Look at Gemini's Language Abilities
#40The Gemini white paper reports higher scores on HumanEval and other tasks. So one of Google lied, this eval has bugs, they borked the deployment is true
Most likely Google has lied. AI playing video games and board games don't translate to real world applications. Many people fail to see that.