Live data from Hacker News

An In-depth Look at Gemini's Language Abilities

arxiv.org

1–10 of 73 posts

Re: An In-depth Look at Gemini's Language Abilities

#3
Has anyone (outside of Google) gotten to play with Gemini Ultra yet? Been hearing a lot about Pro, but I'd be interested in seeing whether Ultra is really close to as capable as they claim.

Also very interesting that Mixtral 8x7B ranks in the same neighborhood as Gemini Pro/GPT 3.5 Turbo/Claude 2.1 while being fully open source and Apache 2.0 licensed.

Re: An In-depth Look at Gemini's Language Abilities

#5
post #3

Has anyone (outside of Google) gotten to play with Gemini Ultra yet? Been hearing a lot about Pro, but I'd be interested in seeing whether Ultra is really close to as capable as they claim. Also very interesting that Mixtral 8x7B ranks in the same neighborhood as Gemini Pro/GPT 3.5 Turbo/Claude 2.1 while being fully open source and Apache 2.0 licensed.

Looks like Microsoft people have had access to it to benchmark promptbase: https://github.com/microsoft/promptbase

Re: An In-depth Look at Gemini's Language Abilities

#6
post #3

Has anyone (outside of Google) gotten to play with Gemini Ultra yet? Been hearing a lot about Pro, but I'd be interested in seeing whether Ultra is really close to as capable as they claim. Also very interesting that Mixtral 8x7B ranks in the same neighborhood as Gemini Pro/GPT 3.5 Turbo/Claude 2.1 while being fully open source and Apache 2.0 licensed.

Looks like Microsoft people have had access to it to benchmark promptbase: https://github.com/microsoft/promptbase

Those benchmarks are from the Gemini Ultra announcement post

Re: An In-depth Look at Gemini's Language Abilities

#7
post #6

Earlier quoted context omitted.

Looks like Microsoft people have had access to it to benchmark promptbase: https://github.com/microsoft/promptbase

Those benchmarks are from the Gemini Ultra announcement post

Oops, sorry, yes you're right, I think. I was skimming the lists of prompts incorrectly.

Re: An In-depth Look at Gemini's Language Abilities

#8

It's incredible how accurate the Chatbot Arena Leaderboard [0] is at predicting model performance compared to benchmarks (which can and are being gamed, see all the 7B models on HF leaderboard) [0]: https://huggingface.co/spaces/lmsys/chatbot-arena-leaderboar...

It's astounding that Mixtral Instruct ties with 3.5-turbo while being ~10x smaller.

Re: An In-depth Look at Gemini's Language Abilities

#9
post #3

Has anyone (outside of Google) gotten to play with Gemini Ultra yet? Been hearing a lot about Pro, but I'd be interested in seeing whether Ultra is really close to as capable as they claim. Also very interesting that Mixtral 8x7B ranks in the same neighborhood as Gemini Pro/GPT 3.5 Turbo/Claude 2.1 while being fully open source and Apache 2.0 licensed.

Mixtral is a mystery to me. How in the world is that team on par with/beating GOOGLE, who presumably have all the resources in the world to throw at this?

Re: An In-depth Look at Gemini's Language Abilities

#10

It's incredible how accurate the Chatbot Arena Leaderboard [0] is at predicting model performance compared to benchmarks (which can and are being gamed, see all the 7B models on HF leaderboard) [0]: https://huggingface.co/spaces/lmsys/chatbot-arena-leaderboar...

The chart on the bottom left corner of that page shows quite how far ahead the various GPT-4 models are compared to everyone else...
Post reply on HN