Live data from Hacker News

An In-depth Look at Gemini's Language Abilities

arxiv.org

41–50 of 73 posts

Re: An In-depth Look at Gemini's Language Abilities

#41
post #35
post #26

Earlier quoted context omitted.

Gemini Ultra is not out yet. With the same logic, you could compare an unreleased Mistral model with Gemini Ultra.

Right. I'm just pointing out that comparing one model with a distilled version of another and then making broad statements about the companies behind them isn't really useful. Surely you could make a comparison of two unreleased models, but it wouldn't be interesting because you don't have any real data (and benchmarks don't really mean anything).

Debating the usefulness of hn commentary is a somewhat philosophical issue, but I think it's entirely fair to draw parallels between what is, not what might be.

Gemini Ultra is self-evidently not ready for production. What the issues are? Who knows, but in a game that as of right now is mostly about reducing the amount of brute force required, something as "simple" as not being efficient enough is actually not something to gloss over. If your engines entire stick is having the greatest graphics but you can't make it run at acceptable fps, well, then it's not actually a usable product.

A LLM that is not actually released could very well be in a comparably dire state and fixing it while also delivering on the promised performance might be entirely non-trivial.

Re: An In-depth Look at Gemini's Language Abilities

#42
post #36

I don't understand why people keep falling for Google's ad campaign. Google have its lead in AI playing video games and board games. It is cool, entertaining and all that jazz. But OpenAI and MS are the real leaders in real AI.

Even if you don't think Google doesn't have a talent or product chops to be leaders in AI, Google can do things cheaper than others because of their infrastructure. When they do release something useful they'll probably be able to offer it free and force it on people on the most visited pages/most used browser. Surprised how many people think having a years head start means OpenAI and Microsoft are going to always be…

It's because no one can see Google doing this without gimping it to prevent it from cannibalizing ads that gives OpenAI staying power.

Re: An In-depth Look at Gemini's Language Abilities

#43
post #3

Has anyone (outside of Google) gotten to play with Gemini Ultra yet? Been hearing a lot about Pro, but I'd be interested in seeing whether Ultra is really close to as capable as they claim. Also very interesting that Mixtral 8x7B ranks in the same neighborhood as Gemini Pro/GPT 3.5 Turbo/Claude 2.1 while being fully open source and Apache 2.0 licensed.

The hosts of the All In podcast have used it, but they're billionaires. They think highly of Ultra. Early on they just talk about the paper that's released, but then they drop that they've used Ultra. https://www.youtube.com/watch?v=IeKUcpU5-Xk&t=3667s

Started watching the video wondering when it would turn into a rant about "wokeness" and "cancel culture", and it happened about 30 seconds in. Glad to see these guys haven't changed.

Re: An In-depth Look at Gemini's Language Abilities

#44
post #26
post #12

Earlier quoted context omitted.

Mixtral is on-par with Gemini Pro, not Gemini Ultra (and even there it is further behind Gemini Pro than Gemini Pro is behind GPT 3.5). But to directly answer your question, they are quite well-funded, having raised over $700mil to date. I definitely wouldn't count them out.

Gemini Ultra is not out yet. With the same logic, you could compare an unreleased Mistral model with Gemini Ultra.

Mistral “Medium” is available (in beta, via API) and should give better results than the “Small” mixtral model.

Re: An In-depth Look at Gemini's Language Abilities

#45

It's incredible how accurate the Chatbot Arena Leaderboard [0] is at predicting model performance compared to benchmarks (which can and are being gamed, see all the 7B models on HF leaderboard) [0]: https://huggingface.co/spaces/lmsys/chatbot-arena-leaderboar...

I wish that Arena included a few more "interesting" models like the new Phi-2 model and the current tinyllama model, which are trying to push the limits on small models. Solar-10.7B is another interesting model that seems to be missing, but I just learned about it yesterday, and it seems to have come out a week ago, so maybe it's too new. Solar supposedly outperforms Mixtral-8x7B with a fraction of the total parameters, although Solar seems optimized for single-turn conversation, so maybe it falls apart over multiple messages (I'm not sure).

Re: An In-depth Look at Gemini's Language Abilities

#46

It's incredible how accurate the Chatbot Arena Leaderboard [0] is at predicting model performance compared to benchmarks (which can and are being gamed, see all the 7B models on HF leaderboard) [0]: https://huggingface.co/spaces/lmsys/chatbot-arena-leaderboar...

I wish that Arena included a few more "interesting" models like the new Phi-2 model and the current tinyllama model, which are trying to push the limits on small models. Solar-10.7B is another interesting model that seems to be missing, but I just learned about it yesterday, and it seems to have come out a week ago, so maybe it's too new. Solar supposedly outperforms Mixtral-8x7B with a fraction of the total paramete…

Solar-10.7B is present in the battle arena but there are probably not enough votes for the ranking.

Re: An In-depth Look at Gemini's Language Abilities

#47
post #21

It's incredible how accurate the Chatbot Arena Leaderboard [0] is at predicting model performance compared to benchmarks (which can and are being gamed, see all the 7B models on HF leaderboard) [0]: https://huggingface.co/spaces/lmsys/chatbot-arena-leaderboar...

It's because it isn't "predicting" anything, but rather aggregating user feedback. That is of course going to be closest to judging the subjective "best" model that pleases most people. It's like saying how can evaluating 5 years of performance at work be better at predicting someone's competency than their SAT scores.

But, what if you could make an SAT that is equivalent to evaluating years of performance at work?

https://huggingface.co/papers/2306.05685

This paper makes the argument that...

"Our results reveal that strong LLM judges like GPT-4 can match both controlled and crowdsourced human preferences well, achieving over 80% agreement, the same level of agreement between humans. Hence, LLM-as-a-judge is a scalable and explainable way to approximate human preferences, which are otherwise very expensive to obtain."

So, the Arena could theoretically be automated and achieve similar outcomes. Or at least, it could quickly determine a predicted-ELO for every model, which would be interesting to compare against the human-rated outcomes.

Re: An In-depth Look at Gemini's Language Abilities

#48
Starling 7b alpha is close to gpt3.5 but it asking list 5 capitals sorted by population size and it will run on like forever generating 100 capitals and explainations and gibrish. Not sure if it’s my prompt, but it sure isn’t close to the intelligence of gpt3.5 when asked the same.

Starling:

```

USER list 5 largest capitals sorted by largest population first

AI Here are some countries and their capitals in alphabetical order Australia: Canberra Belgium: Brussels Canada: Ottawa Finland: Helsinki Italy: Rome France: Paris Germany: Berlin Greece: Athens Ireland: Dublin …(lists 50 others)

Gpt3.5

list 5 largest capitals sorted by largest population first

ChatGPT Certainly! To list the 5 largest capital cities by population, I'll provide you with the information as of my knowledge cutoff date in 2022. Please note that population figures may change over time, and it's a good idea to verify the latest data from reliable sources if needed. Here are the 5 largest capital cities sorted by population, starting with the largest:

Beijing, China: Beijing is the capital of China and has a population of over 21 million people as of 2022. New Delhi, India: New Delhi is the capital of India and is part of the larger National Capital Territory of Delhi. It has a population of approximately 16.8 million people as of 2022. Tokyo, Japan: Tokyo is the capital of Japan, and while the city itself is not officially designated as the capital, it serves as the de facto capital. The Tokyo metropolitan area has a population of around 14 million people. Moscow, Russia: Moscow is the capital of Russia and has a population of over 12 million people. Dhaka, Bangladesh: Dhaka is the capital of Bangladesh and has a population of approximately 8.9 million people. ```

Re: An In-depth Look at Gemini's Language Abilities

#49
post #12
post #9

Earlier quoted context omitted.

Mixtral is a mystery to me. How in the world is that team on par with/beating GOOGLE, who presumably have all the resources in the world to throw at this?

Mixtral is on-par with Gemini Pro, not Gemini Ultra (and even there it is further behind Gemini Pro than Gemini Pro is behind GPT 3.5). But to directly answer your question, they are quite well-funded, having raised over $700mil to date. I definitely wouldn't count them out.

Mixtral ranks higher than Gemini Pro on the (subjective) Chatbot Arena Leaderboard: https://huggingface.co/spaces/lmsys/chatbot-arena-leaderboar...

Where are you seeing that it is "further behind Gemini Pro than Gemini Pro is behind GPT 3.5"?

Re: An In-depth Look at Gemini's Language Abilities

#50
post #9

Earlier quoted context omitted.

Mixtral is a mystery to me. How in the world is that team on par with/beating GOOGLE, who presumably have all the resources in the world to throw at this?

There's a survivorship bias going on here. You've never heard of the thousands of teams out there that are Mistral's size but AREN'T getting results that compete on the global stage, but they do exist. But you've heard of Google, whether they're getting it right or not.

"Thousands of teams" is a vast exaggeration. A tiny handful of companies out there have received funding to the tune of a billion dollars for model training like Mixtral. All of them have researchers with loaded resumes, and most are producing stuff of value. The thousands of other startups in the ecosystem are then taking these APIs and adding trivial abstractions on top.
Post reply on HN