Damn. These things aren't AGI... but I don't care. Luna is good enough for me to give a parser spec and have it write one.
Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
11–20 of 43 posts
Re: Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
#12Re: Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
#13Re: Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
#14AI should have started with science from the beginning, not after 4 years.
I am building on top of it with agents to improve scientific workflows.
Re: Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
#15The fact that opus 5 is outperforming fable is odd to me From personal experience, opus 5 feels net inferior to fable on almost every aspect (for coding tasks)
Since this kind of thing is always down to harness, project-type, and other structural constraints, of course your mileage may vary. Fable is probably great for pen-testing, or as a decision-making kernel of other kinds of applications, and way better than Opus at those things. Probably fine for code-review or changing a codebase of a few thousand lines in any language. Actually building that codebase or changing an even bigger one? Woof.
How this fits in with science? IDK but I bet other existing causal reasoning benchmarks might tell the whole story and this is back to stability again. Sometimes having a smart idea is really important! But more often it's important to just not forget what you were doing. What was I talking about? Oh look a squirrel
Re: Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
#16Evals on actual research workflows is the right direction, most agent benches are toy tasks.
Re: Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
#17Not surprised to see Claude significantly higher in scientific intelligence than Sol. You can tell that Claude really does grasp a wide array of highly specific scientific and mathematical nuances... where's codex is just basically for coding and that's it. That's the feel I get from the both of them anyways and I've used both on the 20x plan for the past week at length. Don't get me wrong, codex is great at finding…
If you have time, can you elaborate or give some examples of mathematical nuances?
I am evaluating Sol and Fable on a fairly large dataset of subtly flawed informal mathematical arguments (task is to identify and name propositions with substantially incorrect proofs in a larger body of text), and Sol is saturating the benchmark, while Fable is below 50% even with the most generous grading.
I don't work in the natural sciences, so I suspect you mean something different by "mathematical nuance".
Re: Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
#18I wonder how long it's going to be before self improvement encompasses hardware and materials science, not just code. It's exciting, soon we'll be able to fully hand off scientific, mathematical, and technical progress over to the machines, and then we can fully lay back.
Re: Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
#19Sad to see no mention of Gemini ...
Re: Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
#20Like you were reading my mind. I was waiting for such benchmark to land. This will improve models for such scientific research workflows. AI should have started with science from the beginning, not after 4 years. I am building on top of it with agents to improve scientific workflows.
DeepMind has entered the chat.