This all feels like a 2024 re-run. Oh, ChatGPT is going to cure cancer? Then find ONE rare cancer and CURE IT. OpenAI has access to the best models and compute - so cure fucking cancer! What the fuck are you waiting for?
Separating signal from noise in coding evaluations
41–50 of 108 posts
Re: Separating signal from noise in coding evaluations
#42In fact, one thing that still bothers me after months is the gpt-5.5 official submission. This task in particular https://www.tbench.ai/leaderboard/terminal-bench/2.0/codex/0...
The task has the following timeouts (https://github.com/harbor-framework/terminal-bench-2/blob/ma...).
[verifier]
timeout_sec = 1200.0
[agent]
timeout_sec = 1200.0
[environment]
build_timeout_sec = 600.0
Which means no agent should take more than 3000 seconds doing it. Two out of five attempts in the link above took well over 3000 seconds (75min and 80 min respectively). Even though they failed, the fact that they ran that long is sus.
Goodhart’s Law at work
Re: Separating signal from noise in coding evaluations
#43Achieving AGI will be more than just passing all benchmarks, it has to account for the unknown problems too.
AGI is a long way off. Unless you’re talking about some unknown-to-me LLM marketing BS which is called “AGI” or something, I guess. Artificial general purpose intelligence is so different to LLMs or image AI that they are completely incomparable, except to say that they are all artificial. AGI will do a lot more than token prediction.
Re: Separating signal from noise in coding evaluations
#44Fundamentally aren’t they concluding that tasks assigned to software developers (human or otherwise) are often incomplete, self contradictory or worse? This is the world in which their tool must play. I’m unsympathetic.
So is the argument that frontier models are not just junior engineers, but first-month interns with no capability of progressing beyond that level?
Re: Separating signal from noise in coding evaluations
#45Re: Separating signal from noise in coding evaluations
#46Give us something that measures a combination of efficiency and intelligence.
I think this would allow for some interesting tactics for smaller models - eg they could do things like computer use to test their results and grind on problems for longer to verify the outputs, whereas larger models may not have budget to self-test.
Re: Separating signal from noise in coding evaluations
#47Fundamentally aren’t they concluding that tasks assigned to software developers (human or otherwise) are often incomplete, self contradictory or worse? This is the world in which their tool must play. I’m unsympathetic.
The more subtle point is that there's a gap between the task and its verification. e.g. if you have an open-ended / under-specified prompt, the verification needs to be able to handle all potential solutions. So you can have a very narrow task prompt that's easy to verify (but likely too simple of a challenge). Or a more realistic task prompt that's much harder to verify. And likely harder to both build the robust ve…
It's not a pipeline, it's an ongoing conversation within any functional team, but this requires buy-in from management, who is often selected for "line must go up this quarter no matter the cost" over "hey, wouldn't it be cool if this company was still a going concern in twenty years?"
Re: Separating signal from noise in coding evaluations
#48I want a new bench - given $100 of api spend, how much can a model accomplish for a suite of benchmark tests? Give us something that measures a combination of efficiency and intelligence. I think this would allow for some interesting tactics for smaller models - eg they could do things like computer use to test their results and grind on problems for longer to verify the outputs, whereas larger models may not have bu…
Toby Ord did what he could with public data and it… doesn’t look great.
Re: Separating signal from noise in coding evaluations
#49I want a new bench - given $100 of api spend, how much can a model accomplish for a suite of benchmark tests? Give us something that measures a combination of efficiency and intelligence. I think this would allow for some interesting tactics for smaller models - eg they could do things like computer use to test their results and grind on problems for longer to verify the outputs, whereas larger models may not have bu…
https://artificialanalysis.ai/?cost=intelligence-vs-cost-per...