Live data from Hacker News

Kagi LLM Benchmarking Project

help.kagi.com

11–15 of 15 posts

Re: Kagi LLM Benchmarking Project

#11
post #8
post #3

This is almost perfect. The gold standard for LLM evaluation would have the following qualities: 1. Categorized (e.g. coding, reasoning, general knowledge) 2. Multimodal (at least text and image) 3. Multiple difficulties (something like "GPT-4 saturates or scores >90%", a la MMLU, "GPT-4 scores 20-80%", and "GPT-4 scores 4. Hidden (under 10% of the dataset publicly available, enough methodological detail to inspire c…

> Hidden It's pretty absurd that it can be a criteria for a good benchmark. You should dismiss any paper (e.g. benchmark result) that isn't repeatable, and by definition closed testing dataset is not repeatable and in fact doesn't provide much insight anyway, you can as well call it arbitrary curated rating, like those that useless journalists do ("top 50 most influential women of all time"). Obviously, this contradi…

One way to handle this might be to have the data hidden, but verifiable in the future. That is: publish a signed hash of the benchmark questions, and every X amount of time swap them out and publish the old ones.

> I don't have a solution. I'm just saying this is complete bullshit, the idea that we are starting to exclaim "yay, hidden data! I can trust that!" just cannot be acceptable.

This feels just unkind. If you acknowledge that it's a hard (impossible?) problem, you should give some leeway to people doing their best until a consensus on the right approach(es) exists.

Re: Kagi LLM Benchmarking Project

#14
post #3

This is almost perfect. The gold standard for LLM evaluation would have the following qualities: 1. Categorized (e.g. coding, reasoning, general knowledge) 2. Multimodal (at least text and image) 3. Multiple difficulties (something like "GPT-4 saturates or scores >90%", a la MMLU, "GPT-4 scores 20-80%", and "GPT-4 scores 4. Hidden (under 10% of the dataset publicly available, enough methodological detail to inspire c…

9
Post reply on HN