Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
21–30 of 43 posts
Re: Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
#22Not surprised to see Claude significantly higher in scientific intelligence than Sol. You can tell that Claude really does grasp a wide array of highly specific scientific and mathematical nuances... where's codex is just basically for coding and that's it. That's the feel I get from the both of them anyways and I've used both on the 20x plan for the past week at length. Don't get me wrong, codex is great at finding…
Re: Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
#23Not surprised to see Claude significantly higher in scientific intelligence than Sol. You can tell that Claude really does grasp a wide array of highly specific scientific and mathematical nuances... where's codex is just basically for coding and that's it. That's the feel I get from the both of them anyways and I've used both on the 20x plan for the past week at length. Don't get me wrong, codex is great at finding…
Re: Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
#24Re: Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
#25Re: Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
#26Software can rely on layers of testing and verification and most code is applying decades-old patterns to a customer's donain and gluing libraries together until they click. That simply don't work when you're on the frontier of something entirely new. There was already a glut of slop science before 2024. We don't need even more science-shaped slop clogging the peer review pipelines. Trust in science is at an all time low. This is a terrible thing to benchmark for and optimize.
Re: Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
#27I wonder how long it's going to be before self improvement encompasses hardware and materials science, not just code. It's exciting, soon we'll be able to fully hand off scientific, mathematical, and technical progress over to the machines, and then we can fully lay back.
I love it when people handwave tech developments of a truly gargantuan scale
Re: Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
#28I worry this doesn’t check correctness. I’ve been finding Claude is lately awful at folllowing instructions, I’ll ask it to implement the algorithm from a paper and it will do something simpler and slower and when challenged do it’s stupid apology thing. It can’t be trusted with anything I’d put in a paper, it lies too much. 4.6 couldn’t do as complex tasks, but it would do what was asked.
Then it's not a valid benchmark. I agree though they're not reliable enough to just put results in a paper.
Re: Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
#29https://github.com/harbor-framework/terminal-bench-science/t...