Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
terminal-bench-science.ai
Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
1–10 of 43 posts
Re: Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
#2Re: Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
#3Re: Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
#4Luna is good enough for me to give a parser spec and have it write one.
Re: Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
#5Evals on actual research workflows is the right direction, most agent benches are toy tasks.
Re: Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
#6Re: Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
#7From personal experience, opus 5 feels net inferior to fable on almost every aspect (for coding tasks)
Re: Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
#8The fact that opus 5 is outperforming fable is odd to me From personal experience, opus 5 feels net inferior to fable on almost every aspect (for coding tasks)
Re: Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
#9You can tell that Claude really does grasp a wide array of highly specific scientific and mathematical nuances... where's codex is just basically for coding and that's it.
That's the feel I get from the both of them anyways and I've used both on the 20x plan for the past week at length.
Don't get me wrong, codex is great at finding bugs and building games. It's great.