Live data from Hacker News

Terminal-Bench-Science: Evaluating AI agents on scientific research workflows

terminal-bench-science.ai

21–30 of 43 posts

Re: Terminal-Bench-Science: Evaluating AI agents on scientific research workflows

#21
I worry this doesn’t check correctness. I’ve been finding Claude is lately awful at folllowing instructions, I’ll ask it to implement the algorithm from a paper and it will do something simpler and slower and when challenged do it’s stupid apology thing. It can’t be trusted with anything I’d put in a paper, it lies too much. 4.6 couldn’t do as complex tasks, but it would do what was asked.

Re: Terminal-Bench-Science: Evaluating AI agents on scientific research workflows

#22

Not surprised to see Claude significantly higher in scientific intelligence than Sol. You can tell that Claude really does grasp a wide array of highly specific scientific and mathematical nuances... where's codex is just basically for coding and that's it. That's the feel I get from the both of them anyways and I've used both on the 20x plan for the past week at length. Don't get me wrong, codex is great at finding…

[deleted]

Re: Terminal-Bench-Science: Evaluating AI agents on scientific research workflows

#23

Not surprised to see Claude significantly higher in scientific intelligence than Sol. You can tell that Claude really does grasp a wide array of highly specific scientific and mathematical nuances... where's codex is just basically for coding and that's it. That's the feel I get from the both of them anyways and I've used both on the 20x plan for the past week at length. Don't get me wrong, codex is great at finding…

So why is Sol the one solving Erdős problems?

Re: Terminal-Bench-Science: Evaluating AI agents on scientific research workflows

#26
No. Please no. I don't want science vibecoded.

Software can rely on layers of testing and verification and most code is applying decades-old patterns to a customer's donain and gluing libraries together until they click. That simply don't work when you're on the frontier of something entirely new. There was already a glut of slop science before 2024. We don't need even more science-shaped slop clogging the peer review pipelines. Trust in science is at an all time low. This is a terrible thing to benchmark for and optimize.

Re: Terminal-Bench-Science: Evaluating AI agents on scientific research workflows

#27

I wonder how long it's going to be before self improvement encompasses hardware and materials science, not just code. It's exciting, soon we'll be able to fully hand off scientific, mathematical, and technical progress over to the machines, and then we can fully lay back.

I love it when people handwave tech developments of a truly gargantuan scale

When ChatGPT4 was released I was like: "shit we're going to have programmable matter way earlier than I thought. Ferrari! Materialize! But in yellow this time!".

Re: Terminal-Bench-Science: Evaluating AI agents on scientific research workflows

#28

I worry this doesn’t check correctness. I’ve been finding Claude is lately awful at folllowing instructions, I’ll ask it to implement the algorithm from a paper and it will do something simpler and slower and when challenged do it’s stupid apology thing. It can’t be trusted with anything I’d put in a paper, it lies too much. 4.6 couldn’t do as complex tasks, but it would do what was asked.

> I worry this doesn’t check correctness

Then it's not a valid benchmark. I agree though they're not reliable enough to just put results in a paper.

Post reply on HN