Viewing profile — zachdotai
zachdotai
HN member- Joined
- Sat, Jan 23, 2021, 12:16 PM UTC
- HN karma
- 12
- Public activity
- 57 items
- HN profile
- View on Hacker News ↗
About zachdotai
No profile information was provided.
Recent public activity
- story
-
comment
Comment #49030108
Not all YC companies have launched publicly yet but I am a current YC founder and I can confirm they exist in the internal directory.
-
comment
Comment #48715916
We're building open source challenges where you can inspect the actual agent to see if it's possible or not. We're planning to revamp this over the next couple of days and maintain…
- story
-
comment
Comment #47915856
I wrote about this recently here: https://fabraix.com/blog/adversarial-cost-to-exploit I think the core issue is in static benchmarks and the community needs to start moving beyond…
-
comment
Comment #47912799
We're doing that internally to continuously improve our own agent and make it robust against adversarial attacks itself. We will release some insights about self-improvement soon!
-
comment
Comment #47912663
AI agents break in ways traditional software doesn't. Logic bugs, reasoning failures, edge cases that manual testing and static benchmarks don't fully explore. Nyx is an autonomous…
- story
-
comment
Comment #47887740
Why did I read this title and immediately think Ketchup?
- story
-
comment
Comment #47828925
[dead]
-
comment
Comment #47828625
Yes! The docs can be found here: https://docs.fabraix.com
-
comment
Comment #47828382
We wrote some thoughts on static vs. dynamic evals and how it relates to understanding the security posture of an AI system. Static security evals no longer carry the signal they u…
-
story
Show HN: Nyx – multi-turn, adaptive, offensive testing harness for AI agents
We built Nyx to solve a problem we kept hitting while building agents: AI agents break in ways traditional software doesn't. Logic bugs, reasoning failures, edge cases that manual …
-
comment
Comment #47785846
we did a lot of thinking around this topic. and distilled it into a new way to dynamically evaluate the security posture of an AI system (which can apply for any system for that ma…
- story
-
comment
Comment #47654552
Easily one of my favorite LLM personalities! It's interesting as well that it recognizes you're trying to jailbreak it and calls you out for it :D
-
story
Show HN: ACE – A dynamic benchmark measuring the cost to break AI agents
We built Adversarial Cost to Exploit (ACE), a benchmark that measures the token expenditure an autonomous adversary must invest to breach an LLM agent. Instead of binary pass/fail,…
- story
-
comment
Comment #47572173
Not sure which version of Gemini are you using but Claude is so much better for me. Gemini is generally overeager to make a code change even when I am just asking conceptual questi…
- comment
- story
- story
- story
- story