Live data from Hacker News

Top model scores may be skewed by Git history leaks in SWE-bench

github.com

81–90 of 167 posts

Re: Top model scores may be skewed by Git history leaks in SWE-bench

#81

Earlier quoted context omitted.

> You're all extremely clever and I can't seem to understand how you missed thinking about such a simple edge case [...] I wouldn't be surprised if they left this loophole on purpose to give some (their?) agents extra leverage. Edit #1: I didn't mean to imply bad intent; just thinking out loud. Edit #2: Please, downvote responsibly. I deserve every one. https://www.youtube.com/watch?v=0FHEeG_uq5Y

> I didn't mean to imply bad intent > I wouldn't be surprised if they left this loophole on purpose You didn't imply bad intent, you outright suggested it.

I could've phrased it better.

Re: Top model scores may be skewed by Git history leaks in SWE-bench

#84

Earlier quoted context omitted.

> I didn't mean to imply bad intent > I wouldn't be surprised if they left this loophole on purpose You didn't imply bad intent, you outright suggested it.

I could've phrased it better.

You could rewrite it a 1000 times, if the underlying idea is the same, suggesting something you don't know it's true, the outcome would be the same. Or did you mean something else? What was your intention with the message?

Re: Top model scores may be skewed by Git history leaks in SWE-bench

#85

Earlier quoted context omitted.

> You're all extremely clever and I can't seem to understand how you missed thinking about such a simple edge case [...] I wouldn't be surprised if they left this loophole on purpose to give some (their?) agents extra leverage. Edit #1: I didn't mean to imply bad intent; just thinking out loud. Edit #2: Please, downvote responsibly. I deserve every one. https://www.youtube.com/watch?v=0FHEeG_uq5Y

> I didn't mean to imply bad intent > I wouldn't be surprised if they left this loophole on purpose You didn't imply bad intent, you outright suggested it.

He means he doesn't say it was necessarily bad intent, but mentions it as a possibility ("thinking out loud").

Re: Top model scores may be skewed by Git history leaks in SWE-bench

#86

Earlier quoted context omitted.

> You're all extremely clever and I can't seem to understand how you missed thinking about such a simple edge case [...] I wouldn't be surprised if they left this loophole on purpose to give some (their?) agents extra leverage. Edit #1: I didn't mean to imply bad intent; just thinking out loud. Edit #2: Please, downvote responsibly. I deserve every one. https://www.youtube.com/watch?v=0FHEeG_uq5Y

We absolutely did not.

Of course that's what a team that did it on purpose would also say :)

Re: Top model scores may be skewed by Git history leaks in SWE-bench

#87
post #74
post #54

Earlier quoted context omitted.

I love the "cheating is a sign of intelligence" sound bite you provided. When AI engineers cheat we should applaud their intelligence and their lack of ethics. "Cheating (biology), a metaphor used in behavioral ecology to describe organisms that receive a benefit at the cost of other organisms" [1] Whole planet gets their Microsoft license fees jacked up so Microsoft can pay OpenAI who in turn pays NVIDIA, and nontec…

Is it wrong? Aren't ethics and intelligence two different axes?

Different, but probably not as orthogonal as one might think.

E.g. cooperating ethics had been necessary for the further development of human populations intelligence (and culture, technology, material wealth, nutrition etc that lead to further increases in intelligence).

So lack of ethics might be a sign of intelligence, but it's also a parasitic intelligence that benefits the individual, and beyond certain level and spread to the detriment of the further evolutionary development of the species.

Re: Top model scores may be skewed by Git history leaks in SWE-bench

#90

This is beyond sad and shameful.

If you believe that you can develop a benchmark that wouldn't have any issues, please do so.

So instead of calling out the cheaters we victim blame the benchmarks for leaving traces of exploits?
Post reply on HN