The most important part is the database that the agent can see and how clean the data is. I pitched a custom enterprise agent to a client thinking it would be maybe 50/50 time on data vs agent tuning, but it's more like 99/1. The alignment process goes very quickly once you have all the fish in exactly one barrel. I think pulling data dynamically from the source systems is where this turns into a game of whack-a-mole…
I have a question: How does connecting agent to db directly work in case of multi tenant system? There is a high chance that agent can snoop into multiple tenants and mess up the responses
Building reliable agentic AI systems
61–70 of 74 posts
Re: Building reliable agentic AI systems
#62Earlier quoted context omitted.
Author here. A couple of things worth clarifying. The 3.1/5.0 score in the Frontiers paper is a user satisfaction rating on feature completeness. Researchers were asked how well the system met all of their needs, including features that simply didn't exist yet at that point. It's a product maturity signal, not an accuracy or reliability number. The paper is also about a year old and the system has moved on significan…
> On hallucinations, I'd push back on the framing a bit. > It's not a perfect answer, but it's a serious one. Thanks Claude!
From the paper, "we collected feedback from 15 to 20 frequent users". Is it 15 or 20?
I lot of interpretive claims, it reads more like a marketing case study.
Re: Building reliable agentic AI systems
#63Re: Building reliable agentic AI systems
#64This is an ongoing working project since 2024, I would like to see some KPI metrics to back off any productivity /job satisfaction improvement in the research department, or what have you, at Bayer.
Monthly average token usage would be another interesting information to read about. Paired with any latency numbers (time to first token, for example).
Re: Building reliable agentic AI systems
#65Re: Building reliable agentic AI systems
#66Re: Building reliable agentic AI systems
#67Seeing this article and seeing the replies. Oof. Maybe Thoughtworks did some good work in traditional software engineering (not sure) - but why would you trust them to touch anything related to LLMs. They don't seem to know what they are doing.
They trade on the brand name and Martin Fowler's reputation, but even in their heyday were considered pedantic architecture astronauts by many of us trying to get shit done.
Re: Building reliable agentic AI systems
#68Not sure how you manage to measure Faithfulness and Answer Relevancy on the live system, without the ground truth.
Good that you have evals in place, but the user satisfaction score might suggest running ablations on the system would be beneficial. I would start by reducing the iterations and unnecessary steps from the agent.
Re: Building reliable agentic AI systems
#69Re: Building reliable agentic AI systems
#70WE'RE WORKING ON AN ARCHITECTURE, YOU CAN OUR EARLY TESTERS IF YOU WANT