Live data from Hacker News

Senior SWE-Bench: open-source benchmark that assesses agents as senior engineers

senior-swe-bench.snorkel.ai

121–130 of 130 posts

Re: Senior SWE-Bench: open-source benchmark that assesses agents as senior engineers

#121
post #20

I saw on Twitter that in an ML course at Tsinghua University, one of the tests asks students to write quizzes that fail the most LLM models as possible. What if we create a benchmark that works like this and assigns ELO scores? Models fight head-to-head by writing a question, a bug, or an incomplete implementation, which the opponent has to answer, fix, or finish.

We tried this, and it works :) https://arxiv.org/abs/2508.06111

You have to be careful about degenerate / duplicate Qs, as a sibling commenter mentioned.

Recently though, we found that reasoning models have trouble making code-output-prediction tasks (the initial family of verifiable tasks we started with) which other reasoning models can't solve.

We started looking into harder / more agentic tasks (e.g. passing tests, using AISI's Inspect framework) but deprioritised.

Re: Senior SWE-Bench: open-source benchmark that assesses agents as senior engineers

#122

Earlier quoted context omitted.

same observation here opus 4.8 (and i dont understand the people defending gpt 5.5 constantly) was significantly mature, it would even push back against anything off putting where as GPT 5.5 will happily agree and do what is asked but I would note that it takes several tries. 4.8 also requires more than one prompt but its output is significantly higher quality and offers more insight Fable 5 is a different beast howe…

For me it's the exact opposite, Anthropics models seem great for "vibe coding" by non engineers. My girlfriend uses Claude and loves it because she doesn't know any of the terminology and Claude happily fills in the gaps. For me, with 20 years experience engineering across the stack for venture backed companies to FAANG, I cannot handle Claude at all, it writes way too much garbage that I never asked for. Codex is li…

i use both and my view is amateurs often side/simp for codex or claude and professionals use both without emotion.

Re: Senior SWE-Bench: open-source benchmark that assesses agents as senior engineers

#123
post #98

[flagged]

Would you please stop posting like this? You've been doing it repeatedly, and it degrades the threads. In fact the majority of your recent comments have been this sort of shallow, dimissive, snarky stuff. That is not what this site is for, and destroys what it is for. If you want to express your substantive points thoughtfully, that of course would be fine. If you'd please review https://news.ycombinator.com/newsguid…

Unfortunately I'm telling the truth. You guys have a serious quality control problem for both the articles and the incentivized comment sections.

It's all political. Millions of dollars to be on the front page of HN.

It all breaks your post rules way more than my comments do.

But this is low hanging fruit so you go for this.

Re: Senior SWE-Bench: open-source benchmark that assesses agents as senior engineers

#124

Earlier quoted context omitted.

For me it's the exact opposite, Anthropics models seem great for "vibe coding" by non engineers. My girlfriend uses Claude and loves it because she doesn't know any of the terminology and Claude happily fills in the gaps. For me, with 20 years experience engineering across the stack for venture backed companies to FAANG, I cannot handle Claude at all, it writes way too much garbage that I never asked for. Codex is li…

i use both and my view is amateurs often side/simp for codex or claude and professionals use both without emotion.

I use both in different scenarios

Re: Senior SWE-Bench: open-source benchmark that assesses agents as senior engineers

#125
post #97

Earlier quoted context omitted.

So software engineering quality is vibes. All coding is vibe coding.

Could you please not post in the flamewar style to HN? We're trying to avoid that here, and we've had to ask you this many times over the years. If you wouldn't mind reviewing https://news.ycombinator.com/newsguidelines.html and taking the intended spirit of the site more to heart, we'd be grateful.

It's not a flamewar style comment, and I think you know that. I was responding to someone's suggestion that software engineering quality is determined "by instinct" - as in, by feel, as in, vibes. It's a pointed argument about conventions of the software industry, made by repeating what the person said using different words that mean the same thing, to try to get them to see how ridiculous their argument is.

You have warned me several times over the years over valid arguments I've made, with no rhyme or reason other than you personally don't agree with my views (or two people downvote it). It would be nice if you didn't throw your weight around so casually, but then again this is your world, not "the community"'s.

Re: Senior SWE-Bench: open-source benchmark that assesses agents as senior engineers

#126
post #97

Earlier quoted context omitted.

Could you please not post in the flamewar style to HN? We're trying to avoid that here, and we've had to ask you this many times over the years. If you wouldn't mind reviewing https://news.ycombinator.com/newsguidelines.html and taking the intended spirit of the site more to heart, we'd be grateful.

It's not a flamewar style comment, and I think you know that. I was responding to someone's suggestion that software engineering quality is determined "by instinct" - as in, by feel, as in, vibes. It's a pointed argument about conventions of the software industry, made by repeating what the person said using different words that mean the same thing, to try to get them to see how ridiculous their argument is. You have…

This is pretty representative of what I'm calling the flamewar style:

> repeating what the person said using different words that mean the same thing, to try to get them to see how ridiculous their argument is

Especially when you do it in a snarky way, without any clarifying information, it comes across as belittling and adds a layer of edginess that isn't needed. There's no need to "get [someone] to see how ridiculous their argument is". It's enough to offer a different view, or explain what you think a more correct view is.

In fact, if the other person senses that you're trying to "get them to see how ridiculous their argument is", or indeed to "get them" to do anything at all, they usually react primarily to that quality, and only secondarily to any argument—because it breaks the implicit contract of curious conversation. This is how threads get into downward spirals.

Re: Senior SWE-Bench: open-source benchmark that assesses agents as senior engineers

#127
post #126

Earlier quoted context omitted.

It's not a flamewar style comment, and I think you know that. I was responding to someone's suggestion that software engineering quality is determined "by instinct" - as in, by feel, as in, vibes. It's a pointed argument about conventions of the software industry, made by repeating what the person said using different words that mean the same thing, to try to get them to see how ridiculous their argument is. You have…

This is pretty representative of what I'm calling the flamewar style: > repeating what the person said using different words that mean the same thing, to try to get them to see how ridiculous their argument is Especially when you do it in a snarky way, without any clarifying information, it comes across as belittling and adds a layer of edginess that isn't needed. There's no need to "get [someone] to see how ridiculo…

> This is pretty representative of what I'm calling the flamewar style

It would be good to have a HN dictionary of terms. The normal internet definition of a flamewar is "a hostile, prolonged argument on the internet that devolves into personal attacks, insults, and aggressive behavior". I said "all coding is vibe coding". Not really a personal attack or aggressive, is it? Glad to know now there's a separate HN definition for this term though. Thank you for correcting my terrible behavior. (That was sarcasm, by the way, not an aggressive attack; I know this can be hard to interpret)

> It's enough to offer a different view, or explain what you think a more correct view is.

I already offered a different view, in detail, and their response was to double down on the opposite of that view. I would be reiterating the same original point. In order to express the inherent flaw in the view, I associated it to another view most people today find abhorrent. It's high risk, but I have nothing to lose, because they already disagreed with my point.

> In fact, if the other person senses that you're trying to "get them to see how ridiculous their argument is", or indeed to "get them" to do anything at all, they usually react primarily to that quality, and only secondarily to any argument

I mostly comment on HN to inform the casual viewer that there is a different view than that of the majority echo chamber. I don't need the person I'm replying to to agree with me. I'm trying to reach the person whose brain isn't yet shut down to logic and reason. I admit sometimes it comes off as dickish, which is regretful, but hopefully still makes my point.

> because it breaks the implicit contract of curious conversation

Almost nobody on the internet has agreed to such a contract. Most people's brains implicitly believe that whatever they already think is correct, unless presented with some irrefutable proof, or argued by a person they implicitly trust. Many people are curious, but not so curious as to want their mind changed by strangers. Which is why people lean on the heuristics of popularity, authority, success, etc to choose what to believe. Most people don't have curious conversations, they have disagreements and agreements.

HN is the epicenter of the tech contrarian, here to disagree with any prevailing wisdom and consider any claims that don't sound like "common sense" as suspect. You've built a wall to keep out intelligence and evidence, and inflate the power of the crowd to make certain opinions seem more valid, depending on who happens to have more traction in the comments. I believe it's a feature built into this platform to increase engagement and controversy to increase site traffic.

Re: Senior SWE-Bench: open-source benchmark that assesses agents as senior engineers

#128
post #126

Earlier quoted context omitted.

This is pretty representative of what I'm calling the flamewar style: > repeating what the person said using different words that mean the same thing, to try to get them to see how ridiculous their argument is Especially when you do it in a snarky way, without any clarifying information, it comes across as belittling and adds a layer of edginess that isn't needed. There's no need to "get [someone] to see how ridiculo…

> This is pretty representative of what I'm calling the flamewar style It would be good to have a HN dictionary of terms. The normal internet definition of a flamewar is "a hostile, prolonged argument on the internet that devolves into personal attacks, insults, and aggressive behavior" . I said "all coding is vibe coding". Not really a personal attack or aggressive, is it? Glad to know now there's a separate HN defi…

You can't be dickish (to use your word) on HN. If you keep doing it, we'll end up banning you the same way we ban other users who won't follow the site guidelines.

Fortunately you can make any of your substantive points without doing that. Since this will also be more effective at your stated goal of informing readers, there is no reason not to.

Re: Senior SWE-Bench: open-source benchmark that assesses agents as senior engineers

#129
post #128

Earlier quoted context omitted.

> This is pretty representative of what I'm calling the flamewar style It would be good to have a HN dictionary of terms. The normal internet definition of a flamewar is "a hostile, prolonged argument on the internet that devolves into personal attacks, insults, and aggressive behavior" . I said "all coding is vibe coding". Not really a personal attack or aggressive, is it? Glad to know now there's a separate HN defi…

You can't be dickish (to use your word) on HN. If you keep doing it, we'll end up banning you the same way we ban other users who won't follow the site guidelines. Fortunately you can make any of your substantive points without doing that. Since this will also be more effective at your stated goal of informing readers, there is no reason not to.

Please delete my account first, then ban me (or delete all content and then keep the user account to ban it).

Re: Senior SWE-Bench: open-source benchmark that assesses agents as senior engineers

#130
post #97

Earlier quoted context omitted.

Could you please not post in the flamewar style to HN? We're trying to avoid that here, and we've had to ask you this many times over the years. If you wouldn't mind reviewing https://news.ycombinator.com/newsguidelines.html and taking the intended spirit of the site more to heart, we'd be grateful.

It's not a flamewar style comment, and I think you know that. I was responding to someone's suggestion that software engineering quality is determined "by instinct" - as in, by feel, as in, vibes. It's a pointed argument about conventions of the software industry, made by repeating what the person said using different words that mean the same thing, to try to get them to see how ridiculous their argument is. You have…

For what it's worth, I actually agree that good software development genuinely is driven by vibes a lot of the time. Sometimes we get to formalize them into laws and rules, but we learn the vibes before we learn the rules; and if we only learn the rules but not the vibes, we don't reliably manage to apply them or overapply them. So I don't view that as an insult, but as a differently-phrased description of my actual view.

So if you want me to view it as ridiculous, you're gonna have to actually engage with the point.

Post reply on HN