Live data from Hacker News

New Gemini model significantly outperforms others on Chatbot Arena (LMSYS)

lmarena.ai

1–10 of 19 posts

Re: New Gemini model significantly outperforms others on Chatbot Arena (LMSYS)

#2
Based on my testing, this model is significantly better than other Gemini models especially with programming/math related tasks. The current Gemini models are pretty useless for anything related to programming/math, but this experiment model puts Gemini ahead of GPT4o, and pretty close to Claude 3.5.

The major problem with Claude 3.5 is you can't have conversation with a large amount of text because you will constantly hit rate limits and it's very annoying.

This model with a 2 million context window is probably the best model right now for programming.

Re: New Gemini model significantly outperforms others on Chatbot Arena (LMSYS)

#4
I feel like it's at the point where I'm not too sure how these rankings impact the my choice of LLM. Every time a new model tops the charts, I'll try them for a bit and go back to claude-3.5-sonnet. Both for coding and day to day questions.

I don't know if I'm just getting used to the claude style of response, or the orangy UI that I kind of find cozy, but I think we need better ways to convey the difference between models.

Re: New Gemini model significantly outperforms others on Chatbot Arena (LMSYS)

#8

I feel like it's at the point where I'm not too sure how these rankings impact the my choice of LLM. Every time a new model tops the charts, I'll try them for a bit and go back to claude-3.5-sonnet. Both for coding and day to day questions. I don't know if I'm just getting used to the claude style of response, or the orangy UI that I kind of find cozy, but I think we need better ways to convey the difference between…

>orangy UI that I kind of find cozy

Yeah it is strangle cozy. I also can't disassociate claude from jean claude van damme and it make giggle thinking he is helping me code.

Re: New Gemini model significantly outperforms others on Chatbot Arena (LMSYS)

#9
post #5

What is the new Gemini model? 1.5-pro-002?

Here is link to this latest one: https://aistudio.google.com/app/prompts/new_chat?model=gemin... 1.5 Pro-002 came out a couple months ago.

Where’s the info on context length etc? Can’t seem to find the official specs page.

Re: New Gemini model significantly outperforms others on Chatbot Arena (LMSYS)

#10
Claude has been my got to, mainly because of the huge context window. But today, that doesn't seem to be the case, or you hit the rate limit pretty quickly and have to wait a whole day.

Google Studio with it's 2M context window + this experimental version could be a good replacement.

Post reply on HN