Viewing profile — draismaa
draismaa
HN member- Joined
- Thu, Apr 18, 2024, 11:05 AM UTC
- HN karma
- 1
- Public activity
- 25 items
- HN profile
- View on Hacker News ↗
About draismaa
No profile information was provided.
Recent public activity
- story
-
comment
Comment #47513992
Every company I work with evals is on the agenda, but often takes so long to get started or it got removed again, just spinned up the first skills in the article, if this works wel…
- story
- story
-
comment
Comment #47350625
We had so many successfull stories with the LangWatch MCP server, an MCP integration that brings agent evaluation infrastructure directly into Claude Code, Cursor, and any MCP-comp…
- story
-
comment
Comment #47200068
Try: https://langwatch.ai/scenario/ Pretty amazing into running simulations at scale
-
comment
Comment #46944962
Clawdbot has been exploding in usage over the past weeks, so we ran a hackday around it at LangWatch and quickly hit a familiar problem: great agent behavior, zero visibility into …
- story
- story
-
comment
Comment #44712123
We open-sourced Scenario, a tiny framework to simulate and test AI agents by using another AI agent — much like how self-driving cars are tested in controlled environments before r…
- story
-
comment
Comment #44604496
Wow, they were really one of the first companies I noticed in this space, always heard very great feedback on the founder when speaking to prospects. All the best in the upcoming j…
-
story
Agent simulations = unit testing for AI?
In traditional software, we write unit tests to catch regressions before they reach users. In AI systems—especially agentic ones that model breaks down. You can test inputs and out…
- story
-
story
We hit a wall testing AI agents, agents simulations works better
We've been working with teams building AI agents (agentic systems, with actual execution) But here's the thing: everyone says “agents are the future,” yet no one really knows how t…
- story
-
comment
Comment #43320325
LLM evaluations are tricky. You can measure accuracy, latency, cost, hallucinations, bias... but what really matters for your app? Instead of relying on generic benchmarks, build y…
- story
-
comment
Comment #42654014
Excited to introduce LangWatch, the tool designed for developers working with LLMs. It allows you to experiment with DSPy optimizers in a simple way and monitor the performance of …
- story
-
comment
Comment #42558605
Awesome to see more opensource tools in this space. In transparency we'r building the oss tool https://github.com/langwatch/langwatch which is tool for tracing and monitoring your …
- comment
- story
- story