Live data from Hacker News

The Leaderboard Illusion

arxiv.org

21–30 of 56 posts

Re: The Leaderboard Illusion

#21
post #7

Earlier quoted context omitted.

Is it? Sounds to me like they run the same experiment many times and keep the "best" results. Which is cheating, or if the same thing is done in biomedical research: research fraud.

Back in the slashdot days I would experiment on changing conversations. This was due to the way SD would rank and show its posts. Anything below a 3 would not change anything. But if you could get in early AND get a +5 on your post you could drive exactly what the conversation was about. Especially if you were engaged a bit and were willing to add a few more posts onto other posts. Basically get in early and get a hi…

Anecdotally, that same technique works on HN.

Re: The Leaderboard Illusion

#22

I think this is a really interesting paper from Cohere, it really feels that at this point in time you can't trust any public benchmark, and you really need your own private evals.

Any tips on coming up with good private evals?

Yes, I wrote something up here on how Andrei Kaparthy evaluated grok 3 -> https://tomhipwell.co/blog/karpathy_s_vibes_check/

I would pick one of two parts of that analysis that are most relevant to you and zoom in. I'd choose something difficult that the model fails at, then look carefully at how the model failures change as you test different model generations.

Re: The Leaderboard Illusion

#23

Earlier quoted context omitted.

Back in the slashdot days I would experiment on changing conversations. This was due to the way SD would rank and show its posts. Anything below a 3 would not change anything. But if you could get in early AND get a +5 on your post you could drive exactly what the conversation was about. Especially if you were engaged a bit and were willing to add a few more posts onto other posts. Basically get in early and get a hi…

Anecdotally, that same technique works on HN.

And Reddit

Re: The Leaderboard Illusion

#24
post #11
post #3

The fact those big LLM developers devote a significant amount of effort to game benchmarks is a big show of confidence that they are making progress towards AGI and will recoup those billions of dollars and man-hours/s

Is this sarcasm? Otherwise I'm not sure how that follows. Seems more reasonable to believe that they're hitting walls and switching to PR and productizing.

Ending a paragraph with "/s" is a moderately common convention for conveying a sarcastic tone through text.

Re: The Leaderboard Illusion

#25

Earlier quoted context omitted.

Back in the slashdot days I would experiment on changing conversations. This was due to the way SD would rank and show its posts. Anything below a 3 would not change anything. But if you could get in early AND get a +5 on your post you could drive exactly what the conversation was about. Especially if you were engaged a bit and were willing to add a few more posts onto other posts. Basically get in early and get a hi…

Anecdotally, that same technique works on HN.

It's intrinsic to any karma system that has a global karma rating, that is, the message has a concrete "karma" value that is the same for all users.

drcongo recently referenced something I sort of wish I had time to build: https://news.ycombinator.com/item?id=43843116 And/or could just go somewhere to use, which is a system where an upvote doesn't mean "everybody needs to see this more" but instead means "I want to see more of this user's comments", and downvotes mean the corresponding opposite. It's more computationally difficult but would create an interestingly different community, especially as further elaborations were built on that. One of the differences would be to mitigate the first-mover advantage in conversations. Instead of it winning you more karma if it appeals to the general public of the relevant site, what it would instead do is expose you to more people. That would produce more upvotes and downvotes in general but wouldn't necessarily impact visibility in the same way.

Re: The Leaderboard Illusion

#26
post #17

Earlier quoted context omitted.

No, even if the benchmarks are private, it's still an issue. Because you can overfit to the benchmark by trying X random variations of the model, and picking the one that performs best on the benchmark It's similar to how I can pass any multiple-choice exam if you let me keep attempting it and tell me my overall score at the end of each attempt - even if you don't tell me which answers were right/wrong

Maybe there should be some rate limiting on it then? I.e., once a month you can benchmark your model. Of course you can submit under different names, but how many company names can someone realistically come up with and register?

So now you want OpenAI to go even wilder in how they name each new model?

Re: The Leaderboard Illusion

#27
post #12

Also, I've been hearing a lot of complaints that Chatbot Arena tends to favor: - Lots of bullet points in every response. - Emoji. ...even at the expense of accurate answers. And I'm beginning to wonder if the sycophantic behavior of recent models ("That's a brilliant and profound idea") is also being driven by Arena scores. Perhaps LLM users actually do want lots of bullets, emoji and fawning praise. But this seems…

> sycophantic behavior of recent models

The funniest example I've seen recently was "Dude. You just said something deep as hell without even flinching. You're 1000% right:"

Re: The Leaderboard Illusion

#28
post #12

Also, I've been hearing a lot of complaints that Chatbot Arena tends to favor: - Lots of bullet points in every response. - Emoji. ...even at the expense of accurate answers. And I'm beginning to wonder if the sycophantic behavior of recent models ("That's a brilliant and profound idea") is also being driven by Arena scores. Perhaps LLM users actually do want lots of bullets, emoji and fawning praise. But this seems…

> sycophantic behavior of recent models The funniest example I've seen recently was "Dude. You just said something deep as hell without even flinching. You're 1000% right:"

This type of response is the quickest way for me to start verbally abusing the LLM.

Re: The Leaderboard Illusion

#30
post #17

Earlier quoted context omitted.

Maybe there should be some rate limiting on it then? I.e., once a month you can benchmark your model. Of course you can submit under different names, but how many company names can someone realistically come up with and register?

So now you want OpenAI to go even wilder in how they name each new model?

1 model per company per month, max.
Post reply on HN