Senior SWE-Bench: open-source benchmark that assesses agents as senior engineers
senior-swe-bench.snorkel.ai
Senior SWE-Bench: open-source benchmark that assesses agents as senior engineers
1–10 of 130 posts
Re: Senior SWE-Bench: open-source benchmark that assesses agents as senior engineers
#2Re: Senior SWE-Bench: open-source benchmark that assesses agents as senior engineers
#3But seriously, as an industry we're terrible at assessing engineering levels, I've worked with "senior engineers" who can't code and I've worked with "junior engineers" who could run rings around them.
Benchmarks like this should be much more precise about what they're actually testing, and what axes they're hard on. We also need to rise above prompts like "you are a senior engineer", it's woo, and it's far better to ask for precise outcomes.
Re: Senior SWE-Bench: open-source benchmark that assesses agents as senior engineers
#4Top solve rate is currently 24% with Opus 4.8... What's a competent human supposed to score?
I wonder if a model could score higher if it had a human at its disposal?
Re: Senior SWE-Bench: open-source benchmark that assesses agents as senior engineers
#5What you really need is an objective benchmark
Re: Senior SWE-Bench: open-source benchmark that assesses agents as senior engineers
#6Why didn't they just make it "Staff SWE-Bench", would be much better smh. /s But seriously, as an industry we're terrible at assessing engineering levels, I've worked with "senior engineers" who can't code and I've worked with "junior engineers" who could run rings around them. Benchmarks like this should be much more precise about what they're actually testing, and what axes they're hard on. We also need to rise abo…
Re: Senior SWE-Bench: open-source benchmark that assesses agents as senior engineers
#7Benchmarks are great, but I feel like there’s a better way this seems quite subjective. What you really need is an objective benchmark
"When are all the software engineers unemployed?"
Re: Senior SWE-Bench: open-source benchmark that assesses agents as senior engineers
#8I don't know what a better approach would look like while still remaining feasible, however this approach of telling a LLM to make a subjective judgement seems fundamentally flawed.
Re: Senior SWE-Bench: open-source benchmark that assesses agents as senior engineers
#9Benchmarks are great, but I feel like there’s a better way this seems quite subjective. What you really need is an objective benchmark