Live data from Hacker News

An In-depth Look at Gemini's Language Abilities

arxiv.org

61–70 of 73 posts

Re: An In-depth Look at Gemini's Language Abilities

#61
post #12

Earlier quoted context omitted.

Mixtral is on-par with Gemini Pro, not Gemini Ultra (and even there it is further behind Gemini Pro than Gemini Pro is behind GPT 3.5). But to directly answer your question, they are quite well-funded, having raised over $700mil to date. I definitely wouldn't count them out.

Mixtral ranks higher than Gemini Pro on the (subjective) Chatbot Arena Leaderboard: https://huggingface.co/spaces/lmsys/chatbot-arena-leaderboar... Where are you seeing that it is "further behind Gemini Pro than Gemini Pro is behind GPT 3.5"?

Presumably in the very article this HN submission is for (https://arxiv.org/pdf/2312.11444.pdf), table 1.

Re: An In-depth Look at Gemini's Language Abilities

#62
post #61

Earlier quoted context omitted.

Mixtral ranks higher than Gemini Pro on the (subjective) Chatbot Arena Leaderboard: https://huggingface.co/spaces/lmsys/chatbot-arena-leaderboar... Where are you seeing that it is "further behind Gemini Pro than Gemini Pro is behind GPT 3.5"?

Presumably in the very article this HN submission is for ( https://arxiv.org/pdf/2312.11444.pdf ), table 1.

Mixtral is missing in half of the benchmarks in that paper. Hardly conclusive. It’s also common knowledge that these benchmarks have a lot of issues[0]. A good litmus test, but not a substitute for actually seeing how the models do in the real world.

On the topic of “hardly conclusive” things, Gemini Pro literally told me just a few minutes ago[1] that the Avatar movies did not have humans in them. There was no funny business in the prompting. At least Mixtral knows that Avatar has humans in it. Most of Gemini Pro’s responses have been fine, but not exceptional.

[0]: one random article talking about these issues: https://www.surgehq.ai//blog/hellaswag-or-hellabad-36-of-thi...

[1]: https://i.imgur.com/En37EJD.png

Re: An In-depth Look at Gemini's Language Abilities

#63
post #9

Earlier quoted context omitted.

Mixtral is a mystery to me. How in the world is that team on par with/beating GOOGLE, who presumably have all the resources in the world to throw at this?

Mistral.AI was founded by three people from Deepmind, they're beating Google because Google no longer has them.

Same as OpenAI, Anthropic, Cohere, Adept and hundreds of other small-mid sized AI startups. When the dust settles and the space gets more mature the exodus from Google Brain/Deepmind over the last few years will be considered this generation's Fairchild moment.

Re: An In-depth Look at Gemini's Language Abilities

#64
Submitters: "Please use the original title, unless it is misleading or linkbait; don't editorialize." - https://news.ycombinator.com/newsguidelines.html

(Submitted title was "Gemini Pro achieves accuracy slightly inferior to GPT 3.5 Turbo".)

If you want to say what you think is important about an article, that's fine, but do it by adding a comment to the thread. Then your view will be on a level playing field with everyone else's: https://hn.algolia.com/?dateRange=all&page=0&prefix=false&so...

Re: An In-depth Look at Gemini's Language Abilities

#65

Earlier quoted context omitted.

But, what if you could make an SAT that is equivalent to evaluating years of performance at work? https://huggingface.co/papers/2306.05685 This paper makes the argument that... "Our results reveal that strong LLM judges like GPT-4 can match both controlled and crowdsourced human preferences well, achieving over 80% agreement, the same level of agreement between humans. Hence, LLM-as-a-judge is a scalable and explaina…

My understanding was that GPT4 evaluation appeared to specifically favour text that GPT4 would generate itself (leading to some bias towards gpt-based fine-tunes), although I can't remember the details

[deleted]

Re: An In-depth Look at Gemini's Language Abilities

#66

Earlier quoted context omitted.

But, what if you could make an SAT that is equivalent to evaluating years of performance at work? https://huggingface.co/papers/2306.05685 This paper makes the argument that... "Our results reveal that strong LLM judges like GPT-4 can match both controlled and crowdsourced human preferences well, achieving over 80% agreement, the same level of agreement between humans. Hence, LLM-as-a-judge is a scalable and explaina…

My understanding was that GPT4 evaluation appeared to specifically favour text that GPT4 would generate itself (leading to some bias towards gpt-based fine-tunes), although I can't remember the details

GPT-4 apparently shows a small bias (10%) towards itself in the paper, and GPT-3.5 apparently did not show any measurable bias towards itself.

Given the possibility of bias, it would make sense to have the judge “recuse” itself from comparisons involving its own output. Between GPT-4, Claude, and soon Gemini Ultra, there should be several strong LLMs to choose from.

I don’t think it would be a replacement for human rating, but it would be interesting to see.

Re: An In-depth Look at Gemini's Language Abilities

#67
post #13

It's incredible how accurate the Chatbot Arena Leaderboard [0] is at predicting model performance compared to benchmarks (which can and are being gamed, see all the 7B models on HF leaderboard) [0]: https://huggingface.co/spaces/lmsys/chatbot-arena-leaderboar...

It's much more accurate than the Open LLM Leaderboard, that's for sure. Human evaluation has always been the gold standard. I just wish we could filter by the votes which were made after only one or two prompts and I hope they don't include the non-blind votes in the results.

Why filter out the votes made after only one or two prompts? A lot of times, a single response is all you need to see.

Do you really need more than this to know which one you’re going to pick? https://i.imgur.com/En37EJD.png

Avatar doesn’t have humans? Seriously?

Re: An In-depth Look at Gemini's Language Abilities

#68
post #13

Earlier quoted context omitted.

It's much more accurate than the Open LLM Leaderboard, that's for sure. Human evaluation has always been the gold standard. I just wish we could filter by the votes which were made after only one or two prompts and I hope they don't include the non-blind votes in the results.

Why filter out the votes made after only one or two prompts? A lot of times, a single response is all you need to see. Do you really need more than this to know which one you’re going to pick? https://i.imgur.com/En37EJD.png Avatar doesn’t have humans? Seriously?

The thought is, the more a person has used a model, the better they are at evaluating whether or not it is truly worse than another. You can't know if a model is better than another with a sample size of one.

Your test isn't checking for instructions, consistency, logic, just one fact which the model you chose may have gotten right by chance. It's fine assuming you only expect the model to fact check and you don't plan to have a conversation, but if you want more than that, it doesn't work very well.

I'm hoping there are votes in there which can reflect those qualities and filtering by conversation length seems like the easiest way to improve the vote quality a bit.

Re: An In-depth Look at Gemini's Language Abilities

#69
post #18
post #13

Earlier quoted context omitted.

It's much more accurate than the Open LLM Leaderboard, that's for sure. Human evaluation has always been the gold standard. I just wish we could filter by the votes which were made after only one or two prompts and I hope they don't include the non-blind votes in the results.

These are the rules of the battle arena: -Ask any question to two anonymous models (e.g., ChatGPT, Claude, Llama) and vote for the better one! -You can continue chatting until you identify a winner. -Vote won’t be counted if model identity is revealed during conversation.

Perfect ty!

Re: An In-depth Look at Gemini's Language Abilities

#70
Does anyone else have the sinking feeling that GPT 4 is as good as things will get for quite a while?

Someone described LLMs as “blurry JPEGs of the Internet”.

In that sense, maybe GPT 4 is as smart as the hive mind of the Internet gets, and newer models just take sharper pictures but of the same subject. Perhaps GPT 4 trained on one of the best subsets available and everything else is going to be worse or the same…

It’s curious that Sam Altman has publicly stated that OpenAI isn’t working on GPT 5. Why not? Is it because they know it’s a pointless exercise with the current training approaches?

Post reply on HN