> I believe eval startups can work when they're targeting safety benchmarks specifically. Are there any examples of successful startups doing this?
In addition to naming one, I'd also be interesting in whether they actually do rigorous work.
The safety research that tends to get headlines is often extremely misleading, usually with directed prompting, or unreported additions to the system prompt specifying model roleplay behavior.
It's also that the price-value frontier is different for different use-cases. For many of the things that I was doing, I could do harness improvements to make DS V4 Flash catch up in performance with GPT-5.5 or Claude Sonnet, but that's just because of the use-case. And if I'm being honest, this kind of eval doesn't need someone else. Claude and I can build a framework on a per-task thing.
I don't love the anthropomorphization of "Claude and I can build this". Using language like that shapes your thinking in exactly the direction that the AI companies are intending. "Claude" is not a person, it makes as much sense as saying "my laptop and I can build this".
"Safety evals are an exception
I believe eval startups can work when they're targeting safety benchmarks specifically. Researchers who want to work on safety evals tend to be ideologically opposed to working on capabilities, which means they don't migrate to post-training or applications due to monetary incentives."
This is quite interesting. Seems more relevant in 2026.
evals are glorified integration tests, would you invest in an integration test startup? absolutely not. I don't get why we are making all of this fuzz around evals
Because what people actually want is a simple harness to test their use cases against all the frontier models and see which is the cheapest/best for the job. It's simple to say but hard to master doing well, and the important thing is that no matter what tool you have the evals don't write themselves.