However, identifying the right metrics and having the necessary test sets will, at times, be challenging.
Why eval startups fail (2025)
51–60 of 66 posts
Re: Why eval startups fail (2025)
#52I haven't looked that hard, but I can't find articles about this type of eval testing, curious to hear if others have approached writing APIs in this way.
Re: Why eval startups fail (2025)
#53Re: Why eval startups fail (2025)
#54"eval startups have a hard time finding customers, because clients have to be technical developers who want to build with APIs, but also not technical enough to run their own evals" To add to this, they have to be developers who aren't already using a fullstack observability solution, since it's fairly straight forward to add the eval startup featureset to an existing observability solution, and easier (plus cost eff…
Re: Why eval startups fail (2025)
#55Earlier quoted context omitted.
It's an important question! If you are paying a lot of money to use AI models, you care that you are using the best for your task. And it turns out that figuring out which AI models is best for your task is not trivial and requires some expertise.
That was too nice of a reply, I apologize. I just can't understand the thought process and that what exactly are we optimizing for? If you are paying a lot of money to use AI models, you already have so much overhead that precise ranking in an eval is not gonna make much difference between equally "frontier" models. Especially since models are sensitive to the input. So the eval is just gonna evaluate the eval with v…
Re: Why eval startups fail (2025)
#56Earlier quoted context omitted.
(Author) It's short for "evaluation", a test for an AI model. Specifically, an AI evaluation comprises (1) a dataset of prompts (as questions / tasks / queries), (2) some way to score model performance on each prompt, like a set of correct answers or a grading rubric that you can use with an LLM autograder, and (3) a metric, such as accuracy¹. (If you're already familiar with the term "benchmark", it's the same thing…
That sounds a bit weak as a startup idea. Hard to productize, hard to scale, etc. It sounds more like consulting.
Re: Why eval startups fail (2025)
#57Re: Why eval startups fail (2025)
#58Are there any examples of successful startups doing this?