Who knew the real revenue unlock wouldn’t be based on how much paranoid red-teaming the model underwent to resist users jailbreaking its ‘alignment’, and instead more on whether the model is post-trained to use ‘sed’ and ‘git’? Poor Gemini Scientist-heavy orgs that want to solve everything in token space may overtook tool use; meanwhile Anthropic has been super focused on MCP, Claude Code etc for over a year
What do you mean by scientist-heavy orgs solving everything in token space? I feel like they use them to make or run tools almost exclusively.
In other words 'just add a calculator tool' is not as sexy research-wise as making the model accurately eyeball arithmetic in its chain of thought. Maybe I'm wrong but that seems to be the case