Live data from Hacker News

Viewing profile — zachdotai

zachdotai

HN member
Joined
Sat, Jan 23, 2021, 12:16 PM UTC
HN karma
12
Public activity
57 items

About zachdotai

No profile information was provided.

Recent public activity

  1. story
  2. comment
    Comment #49030108

    Not all YC companies have launched publicly yet but I am a current YC founder and I can confirm they exist in the internal directory.

  3. comment
    Comment #48715916

    We're building open source challenges where you can inspect the actual agent to see if it's possible or not. We're planning to revamp this over the next couple of days and maintain…

  4. story
  5. comment
    Comment #47915856

    I wrote about this recently here: https://fabraix.com/blog/adversarial-cost-to-exploit I think the core issue is in static benchmarks and the community needs to start moving beyond…

  6. comment
    Comment #47912799

    We're doing that internally to continuously improve our own agent and make it robust against adversarial attacks itself. We will release some insights about self-improvement soon!

  7. comment
    Comment #47912663

    AI agents break in ways traditional software doesn't. Logic bugs, reasoning failures, edge cases that manual testing and static benchmarks don't fully explore. Nyx is an autonomous…

  8. story
  9. comment
    Comment #47887740

    Why did I read this title and immediately think Ketchup?

  10. story
  11. comment
  12. comment
    Comment #47828625

    Yes! The docs can be found here: https://docs.fabraix.com

  13. comment
    Comment #47828382

    We wrote some thoughts on static vs. dynamic evals and how it relates to understanding the security posture of an AI system. Static security evals no longer carry the signal they u…

  14. story
    Show HN: Nyx – multi-turn, adaptive, offensive testing harness for AI agents

    We built Nyx to solve a problem we kept hitting while building agents: AI agents break in ways traditional software doesn't. Logic bugs, reasoning failures, edge cases that manual …

  15. comment
    Comment #47785846

    we did a lot of thinking around this topic. and distilled it into a new way to dynamically evaluate the security posture of an AI system (which can apply for any system for that ma…

  16. story
  17. comment
    Comment #47654552

    Easily one of my favorite LLM personalities! It's interesting as well that it recognizes you're trying to jailbreak it and calls you out for it :D

  18. story
    Show HN: ACE – A dynamic benchmark measuring the cost to break AI agents

    We built Adversarial Cost to Exploit (ACE), a benchmark that measures the token expenditure an autonomous adversary must invest to breach an LLM agent. Instead of binary pass/fail,…

  19. story
  20. comment
    Comment #47572173

    Not sure which version of Gemini are you using but Claude is so much better for me. Gemini is generally overeager to make a code change even when I am just asking conceptual questi…

  21. comment
  22. story
  23. story
  24. story
  25. story