Online vs. Offline AI Evals: When to Use Each
1–4 of 4 posts
Re: Online vs. Offline AI Evals: When to Use Each
#2Maybe-a-little-hotter-take: you should be using production data for evaluation rather than synthetic events.
Enter Online and offline evals; one measures performance in real time. The other, off the execution path, either before deployment, or after the agent run has happened. When to use both is explained in our latest guide!
Re: Online vs. Offline AI Evals: When to Use Each
#3Not-so-hot take: you need to measure your agent's performance in production. Maybe-a-little-hotter-take: you should be using production data for evaluation rather than synthetic events. Enter Online and offline evals; one measures performance in real time. The other, off the execution path, either before deployment, or after the agent run has happened. When to use both is explained in our latest guide!
I don't know if this is the end state, but I think this is pretty close to what I was seeing a year ago and what we will be seeing next year. Yea, maybe we use agents to do the looking through but you will still be shepherding that.
Re: Online vs. Offline AI Evals: When to Use Each
#4When I saw "load-bearing" I remembered the other discussion in the frontpage explaining how to prevent Claude from saying that so often.