Seeing this article and seeing the replies. Oof. Maybe Thoughtworks did some good work in traditional software engineering (not sure) - but why would you trust them to touch anything related to LLMs. They don't seem to know what they are doing.
Building reliable agentic AI systems
51–60 of 74 posts
Re: Building reliable agentic AI systems
#52Most important piece of information is in the linked Frontiers article: However, the overall capability of the chatbot to fully meet user needs received a lower average score (3.1/5.0), highlighting the need for further improvements. Also there is still the problem of hallucinations, as we see in the „Evaluation“ paragraph: Live traffic evaluations are essential for monitoring system behavior, identifying potential i…
Author here. A couple of things worth clarifying. The 3.1/5.0 score in the Frontiers paper is a user satisfaction rating on feature completeness. Researchers were asked how well the system met all of their needs, including features that simply didn't exist yet at that point. It's a product maturity signal, not an accuracy or reliability number. The paper is also about a year old and the system has moved on significan…
Re: Building reliable agentic AI systems
#53Seeing this article and seeing the replies. Oof. Maybe Thoughtworks did some good work in traditional software engineering (not sure) - but why would you trust them to touch anything related to LLMs. They don't seem to know what they are doing.
Re: Building reliable agentic AI systems
#54Earlier quoted context omitted.
It would not, and you would know that if you actually evaluated the results.
I have gone through this process and evaluated the results. Maybe you're referring to their comment as written, but going through what OC described + handholding leads to very good results in my experience.
Re: Building reliable agentic AI systems
#55Most important piece of information is in the linked Frontiers article: However, the overall capability of the chatbot to fully meet user needs received a lower average score (3.1/5.0), highlighting the need for further improvements. Also there is still the problem of hallucinations, as we see in the „Evaluation“ paragraph: Live traffic evaluations are essential for monitoring system behavior, identifying potential i…
Author here. A couple of things worth clarifying. The 3.1/5.0 score in the Frontiers paper is a user satisfaction rating on feature completeness. Researchers were asked how well the system met all of their needs, including features that simply didn't exist yet at that point. It's a product maturity signal, not an accuracy or reliability number. The paper is also about a year old and the system has moved on significan…
Thanks Claude!
Re: Building reliable agentic AI systems
#56Re: Building reliable agentic AI systems
#57[flagged]
Re: Building reliable agentic AI systems
#58Earlier quoted context omitted.
I have gone through this process and evaluated the results. Maybe you're referring to their comment as written, but going through what OC described + handholding leads to very good results in my experience.
"very good" 99 percent of time and hallucinating 1 percent makes the "very good" part untrustworthy.
I'll take the opportunity to note that if you're running solid evals, you'll have data to back the efficacy of your system. If you are seeing a hallucination rate of 1%, then you certainly should be working on your harness/toolset/context/prompting etc.
Saying "1% hallucination rate..." is akin to saying "30,000mi lifespan for [modern japanese make engine]". Something is wrong.
Re: Building reliable agentic AI systems
#59Most important piece of information is in the linked Frontiers article: However, the overall capability of the chatbot to fully meet user needs received a lower average score (3.1/5.0), highlighting the need for further improvements. Also there is still the problem of hallucinations, as we see in the „Evaluation“ paragraph: Live traffic evaluations are essential for monitoring system behavior, identifying potential i…
Author here. A couple of things worth clarifying. The 3.1/5.0 score in the Frontiers paper is a user satisfaction rating on feature completeness. Researchers were asked how well the system met all of their needs, including features that simply didn't exist yet at that point. It's a product maturity signal, not an accuracy or reliability number. The paper is also about a year old and the system has moved on significan…
Thank you for answering the commenter’s question, your answer was informative and helpful.
However it currently reads like you generated it. Which is a surprise, when you’re trying to foster trust in the quality of your process.