Live data from Hacker News

Viewing profile — draismaa

draismaa

HN member
Joined
Thu, Apr 18, 2024, 11:05 AM UTC
HN karma
1
Public activity
25 items

About draismaa

No profile information was provided.

Recent public activity

  1. story
  2. comment
    Comment #47513992

    Every company I work with evals is on the agenda, but often takes so long to get started or it got removed again, just spinned up the first skills in the article, if this works wel…

  3. story
  4. story
  5. comment
    Comment #47350625

    We had so many successfull stories with the LangWatch MCP server, an MCP integration that brings agent evaluation infrastructure directly into Claude Code, Cursor, and any MCP-comp…

  6. story
  7. comment
    Comment #47200068

    Try: https://langwatch.ai/scenario/ Pretty amazing into running simulations at scale

  8. comment
    Comment #46944962

    Clawdbot has been exploding in usage over the past weeks, so we ran a hackday around it at LangWatch and quickly hit a familiar problem: great agent behavior, zero visibility into …

  9. story
  10. story
  11. comment
    Comment #44712123

    We open-sourced Scenario, a tiny framework to simulate and test AI agents by using another AI agent — much like how self-driving cars are tested in controlled environments before r…

  12. story
  13. comment
    Comment #44604496

    Wow, they were really one of the first companies I noticed in this space, always heard very great feedback on the founder when speaking to prospects. All the best in the upcoming j…

  14. story
    Agent simulations = unit testing for AI?

    In traditional software, we write unit tests to catch regressions before they reach users. In AI systems—especially agentic ones that model breaks down. You can test inputs and out…

  15. story
  16. story
    We hit a wall testing AI agents, agents simulations works better

    We've been working with teams building AI agents (agentic systems, with actual execution) But here's the thing: everyone says “agents are the future,” yet no one really knows how t…

  17. story
  18. comment
    Comment #43320325

    LLM evaluations are tricky. You can measure accuracy, latency, cost, hallucinations, bias... but what really matters for your app? Instead of relying on generic benchmarks, build y…

  19. story
  20. comment
    Comment #42654014

    Excited to introduce LangWatch, the tool designed for developers working with LLMs. It allows you to experiment with DSPy optimizers in a simple way and monitor the performance of …

  21. story
  22. comment
    Comment #42558605

    Awesome to see more opensource tools in this space. In transparency we'r building the oss tool https://github.com/langwatch/langwatch which is tool for tracing and monitoring your …

  23. comment
  24. story
  25. story