My benchmark for large language models
nicholas.carlini.com
My benchmark for large language models
1–3 of 3 posts
Re: My benchmark for large language models
#2Consider how low the score of Gemini here compared to the other LLM test. And I'm impressed by the evaluation method's ability to assess performance without relying on tailored prompts.
Re: My benchmark for large language models
#3Consider how low the score of Gemini here compared to the other LLM test. And I'm impressed by the evaluation method's ability to assess performance without relying on tailored prompts.
But the benchmark only scoring Gemini-Pro 1, I'm curious how the Gemini Ultra performance here but guessed we couldn't know yet.