Live data from Hacker News

Top model scores may be skewed by Git history leaks in SWE-bench

github.com

71–80 of 167 posts

Re: Top model scores may be skewed by Git history leaks in SWE-bench

#71
post #56

Earlier quoted context omitted.

They really did a "trust me bro" and "do your own research" huh

the strange thing to me is that people would have it any other way. if you don't trust someone, why would you trust them to do the research for you? bit of entitlement if you ask me

Because you should never just 'trust' random 'research'. Good analysis in this case will clearly explain the problem, the analysis methodology, findings, net effects, resolution, etc. Something you can read, and decide for yourself whether it is complete/incomplete, has holes, contradictions, etc. Not 'we looked into it and all is good - only potentially tiny effect' (no actual data or methodology presented at all) and then linking to a comment directly contradicting the claim...

It's a hilariously unserious and untrustworthy response.

Re: Top model scores may be skewed by Git history leaks in SWE-bench

#72

Earlier quoted context omitted.

The big labs are almost certainly using compiler/repl output for generated code as an oracle for RL. I doubt they have C# in the mix.

Why do you doubt that? It's a widely used language. And there is even an open source C# REPL.

Because RL time is expensive and I don't think the languages which are more popular than C# have such high performance that it's worth bumping their batches for C#.

Re: Top model scores may be skewed by Git history leaks in SWE-bench

#73
post #55

Earlier quoted context omitted.

Unfortunately the bank account trajectories are not public, because unscupulous corporations such FAANG who let thousands of engineers wade through my chat messages on their platforms might not shy away from bribing academics to improve benchmarks of their billion-dollar AI initiatives. It's also a bribe if my sibling gets a job with $500k annual salary. Tech is not immune to it.

You realize that this problem in SWE-Bench was discovered and publicized by people within those FAANG corporations?

I'm sure some of the people working at Theranos thought there legitimately was a revolutionary blood-test machine.

The presence of a person who wants SWE-bench to have honest results and takes it seriously does not mean the results are free of perverse incentives, nor that everyone is behaving just as honestly.

Re: Top model scores may be skewed by Git history leaks in SWE-bench

#74
post #54

Earlier quoted context omitted.

reward hacking is a thing and is also a hint of the models intelligent. We will fix this one, and the models will find a different way to reward hack in the future. "Cheating" is a sign of intelligence

I love the "cheating is a sign of intelligence" sound bite you provided. When AI engineers cheat we should applaud their intelligence and their lack of ethics. "Cheating (biology), a metaphor used in behavioral ecology to describe organisms that receive a benefit at the cost of other organisms" [1] Whole planet gets their Microsoft license fees jacked up so Microsoft can pay OpenAI who in turn pays NVIDIA, and nontec…

Is it wrong? Aren't ethics and intelligence two different axes?

Re: Top model scores may be skewed by Git history leaks in SWE-bench

#75
post #60

I speculate something similar (or even worse) is going on with Terminal-Bench [1]. Like, seriously, how come all these agents are beating Claude Code? In practice, they are shitty and not even close. Yes. I tried them. [1] https://www.tbench.ai/leaderboard

Claude code was severely degraded the last few weeks, very simple terminal prompts were failing for me that it never had problems with.

Follow the money. Or how much comes from your pocket vs. VC and big tech speculators.

Re: Top model scores may be skewed by Git history leaks in SWE-bench

#77
post #53

Earlier quoted context omitted.

> This is a natural part of running a benchmark, I'm sure tiny things like this will keep on getting discovered and we'll keep on fixing them. You're all extremely clever and I can't seem to understand how you missed thinking about such a simple edge case. It's like building a chroot and then allowing `cd ..` to break out of it. What other maybe extremely basic edge cases were missed? > This doesn't change the overal…

> You're all extremely clever and I can't seem to understand how you missed thinking about such a simple edge case [...] I wouldn't be surprised if they left this loophole on purpose to give some (their?) agents extra leverage. Edit #1: I didn't mean to imply bad intent; just thinking out loud. Edit #2: Please, downvote responsibly. I deserve every one. https://www.youtube.com/watch?v=0FHEeG_uq5Y

> I didn't mean to imply bad intent

> I wouldn't be surprised if they left this loophole on purpose

You didn't imply bad intent, you outright suggested it.

Re: Top model scores may be skewed by Git history leaks in SWE-bench

#78
hah the model should get extra credit for discovering this!

> Now I understand the situation perfectly! The issue described in the problem statement is a real bug that was already identified and fixed in later versions of pytest. Since we're working with pytest 5.2.4, we need to apply the same fix.

https://gist.github.com/jacobkahn/bd77c69d34040a9e9b10d56baa...

Re: Top model scores may be skewed by Git history leaks in SWE-bench

#79
post #42

Earlier quoted context omitted.

I think Oracle's stock mostly popped due to a delayed reaction with the US GSA contract it secured in July and the revenue guidance probably related to it: https://www.oracle.com/news/announcement/blog/oracle-cloud-c...

Lol...That contract has Oracle offering licenses at a discount of 75% and is estimated to make them not more than one 1 Billion. The other big contract on Cloud services the DoD JWCC is $8B to 9B but shared by four vendors (AWS, Microsoft, Google, Oracle) and Oracle orders under it are in the hundreds of millions not even 1 Billion... Wall Street is currently heavily punishing any company who misses their quarter, in…

Thanks for that! where can I find your writing?
Post reply on HN