Benchmarks are great, but I feel like there’s a better way this seems quite subjective. What you really need is an objective benchmark
> What you really need is an objective benchmark "When are all the software engineers unemployed?"
Senior SWE-Bench: open-source benchmark that assesses agents as senior engineers
11–20 of 130 posts
Re: Senior SWE-Bench: open-source benchmark that assesses agents as senior engineers
#12Re: Senior SWE-Bench: open-source benchmark that assesses agents as senior engineers
#13Why didn't they just make it "Staff SWE-Bench", would be much better smh. /s But seriously, as an industry we're terrible at assessing engineering levels, I've worked with "senior engineers" who can't code and I've worked with "junior engineers" who could run rings around them. Benchmarks like this should be much more precise about what they're actually testing, and what axes they're hard on. We also need to rise abo…
As someone who's trying to get better assessments, I'm struggling to come up with objective coding tasks that evaluates all aspects of real life like planning, design choices, problem solving and context usage. From your experience with humans, Do you have any recommendations on what could be effective in measuring it?
Re: Senior SWE-Bench: open-source benchmark that assesses agents as senior engineers
#14Why didn't they just make it "Staff SWE-Bench", would be much better smh. /s But seriously, as an industry we're terrible at assessing engineering levels, I've worked with "senior engineers" who can't code and I've worked with "junior engineers" who could run rings around them. Benchmarks like this should be much more precise about what they're actually testing, and what axes they're hard on. We also need to rise abo…
Re: Senior SWE-Bench: open-source benchmark that assesses agents as senior engineers
#15Re: Senior SWE-Bench: open-source benchmark that assesses agents as senior engineers
#16Benchmarks are great, but I feel like there’s a better way this seems quite subjective. What you really need is an objective benchmark
I actually really like subjective benchmarks, so long as it's a human (ideally me) grading the results. LLM as judge never made much sense.
Re: Senior SWE-Bench: open-source benchmark that assesses agents as senior engineers
#17[flagged]
Re: Senior SWE-Bench: open-source benchmark that assesses agents as senior engineers
#18fable 5?
Re: Senior SWE-Bench: open-source benchmark that assesses agents as senior engineers
#19But it's a lie. Nobody's paying you to make paintings. They're paying you to build machines. The comparison between "making working software" with "taste" always devolves into bikeshedding and subjective opinionism, uses subjective human feelings to describe what should be objective and functional, isn't rooted in scientific rigor, and detracts from the real purpose of the thing. The work doesn't actually get better by trying to apply artistic principles to engineering. It just feels better for the people making it.
Once you make the machine work, then you can go about gilding the lily. But this is unromantic, unsatisfying, boring. Since the inmates run this particular asylum, we end up with a benchmark that tries to accurately mimic the human ego as applied to software design. Thus the new Gods create their digital Adams and Eves in their image.
Re: Senior SWE-Bench: open-source benchmark that assesses agents as senior engineers
#20What if we create a benchmark that works like this and assigns ELO scores? Models fight head-to-head by writing a question, a bug, or an incomplete implementation, which the opponent has to answer, fix, or finish.