Earlier quoted context omitted.
Mixtral is on-par with Gemini Pro, not Gemini Ultra (and even there it is further behind Gemini Pro than Gemini Pro is behind GPT 3.5). But to directly answer your question, they are quite well-funded, having raised over $700mil to date. I definitely wouldn't count them out.
Mixtral ranks higher than Gemini Pro on the (subjective) Chatbot Arena Leaderboard: https://huggingface.co/spaces/lmsys/chatbot-arena-leaderboar... Where are you seeing that it is "further behind Gemini Pro than Gemini Pro is behind GPT 3.5"?
An In-depth Look at Gemini's Language Abilities
61–70 of 73 posts
Re: An In-depth Look at Gemini's Language Abilities
#62Earlier quoted context omitted.
Mixtral ranks higher than Gemini Pro on the (subjective) Chatbot Arena Leaderboard: https://huggingface.co/spaces/lmsys/chatbot-arena-leaderboar... Where are you seeing that it is "further behind Gemini Pro than Gemini Pro is behind GPT 3.5"?
Presumably in the very article this HN submission is for ( https://arxiv.org/pdf/2312.11444.pdf ), table 1.
On the topic of “hardly conclusive” things, Gemini Pro literally told me just a few minutes ago[1] that the Avatar movies did not have humans in them. There was no funny business in the prompting. At least Mixtral knows that Avatar has humans in it. Most of Gemini Pro’s responses have been fine, but not exceptional.
[0]: one random article talking about these issues: https://www.surgehq.ai//blog/hellaswag-or-hellabad-36-of-thi...
Re: An In-depth Look at Gemini's Language Abilities
#63Earlier quoted context omitted.
Mixtral is a mystery to me. How in the world is that team on par with/beating GOOGLE, who presumably have all the resources in the world to throw at this?
Mistral.AI was founded by three people from Deepmind, they're beating Google because Google no longer has them.
Re: An In-depth Look at Gemini's Language Abilities
#64(Submitted title was "Gemini Pro achieves accuracy slightly inferior to GPT 3.5 Turbo".)
If you want to say what you think is important about an article, that's fine, but do it by adding a comment to the thread. Then your view will be on a level playing field with everyone else's: https://hn.algolia.com/?dateRange=all&page=0&prefix=false&so...
Re: An In-depth Look at Gemini's Language Abilities
#65Earlier quoted context omitted.
But, what if you could make an SAT that is equivalent to evaluating years of performance at work? https://huggingface.co/papers/2306.05685 This paper makes the argument that... "Our results reveal that strong LLM judges like GPT-4 can match both controlled and crowdsourced human preferences well, achieving over 80% agreement, the same level of agreement between humans. Hence, LLM-as-a-judge is a scalable and explaina…
My understanding was that GPT4 evaluation appeared to specifically favour text that GPT4 would generate itself (leading to some bias towards gpt-based fine-tunes), although I can't remember the details
Re: An In-depth Look at Gemini's Language Abilities
#66Earlier quoted context omitted.
But, what if you could make an SAT that is equivalent to evaluating years of performance at work? https://huggingface.co/papers/2306.05685 This paper makes the argument that... "Our results reveal that strong LLM judges like GPT-4 can match both controlled and crowdsourced human preferences well, achieving over 80% agreement, the same level of agreement between humans. Hence, LLM-as-a-judge is a scalable and explaina…
My understanding was that GPT4 evaluation appeared to specifically favour text that GPT4 would generate itself (leading to some bias towards gpt-based fine-tunes), although I can't remember the details
Given the possibility of bias, it would make sense to have the judge “recuse” itself from comparisons involving its own output. Between GPT-4, Claude, and soon Gemini Ultra, there should be several strong LLMs to choose from.
I don’t think it would be a replacement for human rating, but it would be interesting to see.
Re: An In-depth Look at Gemini's Language Abilities
#67It's incredible how accurate the Chatbot Arena Leaderboard [0] is at predicting model performance compared to benchmarks (which can and are being gamed, see all the 7B models on HF leaderboard) [0]: https://huggingface.co/spaces/lmsys/chatbot-arena-leaderboar...
It's much more accurate than the Open LLM Leaderboard, that's for sure. Human evaluation has always been the gold standard. I just wish we could filter by the votes which were made after only one or two prompts and I hope they don't include the non-blind votes in the results.
Do you really need more than this to know which one you’re going to pick? https://i.imgur.com/En37EJD.png
Avatar doesn’t have humans? Seriously?
Re: An In-depth Look at Gemini's Language Abilities
#68Earlier quoted context omitted.
It's much more accurate than the Open LLM Leaderboard, that's for sure. Human evaluation has always been the gold standard. I just wish we could filter by the votes which were made after only one or two prompts and I hope they don't include the non-blind votes in the results.
Why filter out the votes made after only one or two prompts? A lot of times, a single response is all you need to see. Do you really need more than this to know which one you’re going to pick? https://i.imgur.com/En37EJD.png Avatar doesn’t have humans? Seriously?
Your test isn't checking for instructions, consistency, logic, just one fact which the model you chose may have gotten right by chance. It's fine assuming you only expect the model to fact check and you don't plan to have a conversation, but if you want more than that, it doesn't work very well.
I'm hoping there are votes in there which can reflect those qualities and filtering by conversation length seems like the easiest way to improve the vote quality a bit.
Re: An In-depth Look at Gemini's Language Abilities
#69Earlier quoted context omitted.
It's much more accurate than the Open LLM Leaderboard, that's for sure. Human evaluation has always been the gold standard. I just wish we could filter by the votes which were made after only one or two prompts and I hope they don't include the non-blind votes in the results.
These are the rules of the battle arena: -Ask any question to two anonymous models (e.g., ChatGPT, Claude, Llama) and vote for the better one! -You can continue chatting until you identify a winner. -Vote won’t be counted if model identity is revealed during conversation.
Re: An In-depth Look at Gemini's Language Abilities
#70Someone described LLMs as “blurry JPEGs of the Internet”.
In that sense, maybe GPT 4 is as smart as the hive mind of the Internet gets, and newer models just take sharper pictures but of the same subject. Perhaps GPT 4 trained on one of the best subsets available and everything else is going to be worse or the same…
It’s curious that Sam Altman has publicly stated that OpenAI isn’t working on GPT 5. Why not? Is it because they know it’s a pointless exercise with the current training approaches?