Live data from Hacker News

The Leaderboard Illusion

arxiv.org

51–56 of 56 posts

Re: The Leaderboard Illusion

#51
I published some notes and opinions on this paper here: https://simonwillison.net/2025/Apr/30/criticism-of-the-chatb...

Short version: the thing I care most about in this paper is that well funded vendors can apparently submit dozens of variations of their models to the leaderboard and then selectively publish the model that did best.

This gives them a huge advantage. I want to know if they did that. A top place model with a footnote saying "they tried 22 variants, most of which scored lower than this one" helps me understand what' going on.

If the top model tried 22 times and scored lower on 21 of those tries, whereas the model in second place only tried once, I'd like to hear about it.

Re: The Leaderboard Illusion

#52

I'm not even following AI model performance testing that closely but I'm hearing increasing reports they're inaccurate due to accidental or intentional test data leaking into training data and other ways of training to the test. Also, ARC AGI reported they've been unable to independently replicate OpenAI's claimed breakthrough score from December. There's just too much money at stake now to not treat all AI model per…

> Also, ARC AGI reported they've been unable to independently replicate OpenAI's claimed breakthrough score from December

Can you elaborate on this? Where did ARC AGI report that? From ARC AGI[0]:

> ARC Prize Foundation was invited by OpenAI to join their “12 Days Of OpenAI.” Here, we shared the results of their first o3 model, o3-preview, on ARC-AGI. It set a new high-water mark for test-time compute, applying near-max resources to the ARC-AGI benchmark.

> We announced that o3-preview (low compute) scored 76% on ARC-AGI-1 Semi Private Eval set and was eligible for our public leaderboard. When we lifted the compute limits, o3-preview (high compute) scored 88%. This was a clear demonstration of what the model could do with unrestricted test-time resources. Both scores were verified to be state of the art.

That makes it sound like ARC AGI were the ones running the original test with o3

What they say they haven't been able to reproduce is o3-preview's performance with the production versions of o3. They attribute this to the production versions being given less compute than the versions they ran in the test

[0] https://arcprize.org/blog/analyzing-o3-with-arc-agi

Re: The Leaderboard Illusion

#53
It’s essentially the pvalue hacking we see in social and biological sciences applied to machine learning field.

Once you set an evaluation metric it ceases to become a useful metric.

Re: The Leaderboard Illusion

#56
post #25

Earlier quoted context omitted.

Anecdotally, that same technique works on HN.

It's intrinsic to any karma system that has a global karma rating, that is, the message has a concrete "karma" value that is the same for all users. drcongo recently referenced something I sort of wish I had time to build: https://news.ycombinator.com/item?id=43843116 And/or could just go somewhere to use, which is a system where an upvote doesn't mean " everybody needs to see this more" but instead means " I want to…

That is an interesting idea. But I suspect it really would still create a moderate first mover advantage in small communities. Early first mover advantage I suspect is decent in any up/down point based system ranking. Would have to run simulations on it. I also suspect what is being described is similar to the way YT works. For example I know they random feed me things. If I click on it and watch the whole vid. Suddenly I get a lot more suggestions from that channel or cohorts to it. But I cant prove that as they are terribly inscrutable on describing what it does (for good reason!).
Post reply on HN