Evaluating 55 LLMs with GPT-4
benchmarks.llmonitor.com
Evaluating 55 LLMs with GPT-4
1–9 of 9 posts
Re: Evaluating 55 LLMs with GPT-4
#2Should be evaluating each prompt multiple times to see how much variance in the scores there are. Even gpt-4 grading gpt-4 should probably be done several times
Re: Evaluating 55 LLMs with GPT-4
#3I always find evals of this flavor offputting given that 3.5 and 4 likely share preference models (or at least feedback data)
Re: Evaluating 55 LLMs with GPT-4
#4Any reason why palm or cohere models are not here ?
Re: Evaluating 55 LLMs with GPT-4
#5Really cool thanks
Re: Evaluating 55 LLMs with GPT-4
#6How is this benchmark not inherently biased towards GPT?
If I did the same sort of thing but used Claude to grade the tests, would I get similar results? Or would that be inherently biased towards Claude scoring high?
Re: Evaluating 55 LLMs with GPT-4
#7GPT-4-0314 is top of the league table (ie. Not the latest version, but the version released in March).
Is this our Concorde moment?
Re: Evaluating 55 LLMs with GPT-4
#8Any reason why palm or cohere models are not here ?
Palm 2 is tied for #10
Re: Evaluating 55 LLMs with GPT-4
#9Why no multi-turn evaluation? A lot of these benchmarks fail to capture the strength of ghost attention used in Llama 2 chat models.