Not surprised to see Claude significantly higher in scientific intelligence than Sol. You can tell that Claude really does grasp a wide array of highly specific scientific and mathematical nuances... where's codex is just basically for coding and that's it. That's the feel I get from the both of them anyways and I've used both on the 20x plan for the past week at length. Don't get me wrong, codex is great at finding…
> You can tell that Claude really does grasp a wide array of highly specific scientific and mathematical nuances... where's codex is just basically for coding and that's it. If you have time, can you elaborate or give some examples of mathematical nuances? I am evaluating Sol and Fable on a fairly large dataset of subtly flawed informal mathematical arguments (task is to identify and name propositions with substantia…
Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
31–40 of 43 posts
Re: Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
#32Re: Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
#33No. Please no. I don't want science vibecoded. Software can rely on layers of testing and verification and most code is applying decades-old patterns to a customer's donain and gluing libraries together until they click. That simply don't work when you're on the frontier of something entirely new. There was already a glut of slop science before 2024. We don't need even more science-shaped slop clogging the peer revie…
Re: Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
#34The tasks are the thing to really look at here: https://github.com/harbor-framework/terminal-bench-science/t...
Re: Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
#35No. Please no. I don't want science vibecoded. Software can rely on layers of testing and verification and most code is applying decades-old patterns to a customer's donain and gluing libraries together until they click. That simply don't work when you're on the frontier of something entirely new. There was already a glut of slop science before 2024. We don't need even more science-shaped slop clogging the peer revie…
Re: Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
#36Earlier quoted context omitted.
> You can tell that Claude really does grasp a wide array of highly specific scientific and mathematical nuances... where's codex is just basically for coding and that's it. If you have time, can you elaborate or give some examples of mathematical nuances? I am evaluating Sol and Fable on a fairly large dataset of subtly flawed informal mathematical arguments (task is to identify and name propositions with substantia…
Is this benchmark public? Anecdotally I have had decent results asking Sol to nitpick my proofs (mostly probability theory but nothing super dense). I have never tried Claude seriously, so I am very curious about what the failures look like with Fable.
Re: Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
#37I worry this doesn’t check correctness. I’ve been finding Claude is lately awful at folllowing instructions, I’ll ask it to implement the algorithm from a paper and it will do something simpler and slower and when challenged do it’s stupid apology thing. It can’t be trusted with anything I’d put in a paper, it lies too much. 4.6 couldn’t do as complex tasks, but it would do what was asked.
Agents are given instructions in markdown format, allowed to read data and libraries sandboxed in a Docker container, and evaluated on deterministic pytests on the outcome. Two things the team aimed to enforce to add tasks we could trust: - Scientific workflows are often simulations that are correct upto numerical tolerances (the scientist decides what's reasonable), so task verifiers' evaluate results of agent-written code within the tolerances. - There could be multiple solution codes to a scientific workflow, and the team tried to ensure the verifier tests accommodate those. Not overfit to the oracle reference code, written by the scientist.
Instruction following is implicitly assumed, if the model gives up and doesn't complete the task it counts as a failure because the verifier tests fail.
Re: Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
#38Not surprised to see Claude significantly higher in scientific intelligence than Sol. You can tell that Claude really does grasp a wide array of highly specific scientific and mathematical nuances... where's codex is just basically for coding and that's it. That's the feel I get from the both of them anyways and I've used both on the 20x plan for the past week at length. Don't get me wrong, codex is great at finding…
So why is Sol the one solving Erdős problems?
For instance, one of the other top comments berates Claude for terrible instruction following with regards to scientific papers, whereas this one is full of praise.
Everyone is just making up their thoughts on these models based on vibes.
Re: Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
#39No. Please no. I don't want science vibecoded. Software can rely on layers of testing and verification and most code is applying decades-old patterns to a customer's donain and gluing libraries together until they click. That simply don't work when you're on the frontier of something entirely new. There was already a glut of slop science before 2024. We don't need even more science-shaped slop clogging the peer revie…
"Vibecoded" science is well on its way to being better than your "real Science™". Withholding the future is not an option.
The signal to noise ratio will be so bad that all of science will suffer.
Re: Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
#40Earlier quoted context omitted.
"Vibecoded" science is well on its way to being better than your "real Science™". Withholding the future is not an option.
I'm sure there will be some golden nuggets in the petabytes of slop that are about to be unleashed. The signal to noise ratio will be so bad that all of science will suffer.