Live data from Hacker News

XBOW, an autonomous penetration tester, has reached the top spot on HackerOne

xbow.com

31–40 of 128 posts

Re: XBOW, an autonomous penetration tester, has reached the top spot on HackerOne

#32

First: > To bridge that gap, we started dogfooding XBOW in public and private bug bounty programs hosted on HackerOne. We treated it like any external researcher would: no shortcuts, no internal knowledge—just XBOW, running on its own. Is it dogfooding if you're not doing it to yourself? I'd considerit dogfooding only if they were flooding themselves in AI generated bug reports, not to other people. They're not the o…

Their success rates on HackerOne seem widely varying.

  22/24 (Valid / Closed) for Walt Disney

  3/43 (Valid / Closed) for AT&T

Re: XBOW, an autonomous penetration tester, has reached the top spot on HackerOne

#34

"XBOW is an enterprise solution. If your company would like a demo, email us at info@xbow.com." Like any "AI" article, this is an ad. If you are willing to tolerate a high false positive rate, you can as well use Rational Purify or various analyzers.

You should come to my upcoming BlackHat talk on how we did this while avoiding false positives :D

https://www.blackhat.com/us-25/briefings/schedule/#ai-agents...

Re: XBOW, an autonomous penetration tester, has reached the top spot on HackerOne

#35
post #20

Earlier quoted context omitted.

Moreover, I don't think XBOW is likely generating the kind of slop beg bounty people generate. There's some serious work behind this.

Still they're sending hundreds of reports that are being refused because they are not following the rules of the bounties. So they better work on that.

If you thought human bounty program participants were generally following the rules, or that programs weren't swamped with slop already... at least these are actually pre-triaged vetted findings.

Re: XBOW, an autonomous penetration tester, has reached the top spot on HackerOne

#36

First: > To bridge that gap, we started dogfooding XBOW in public and private bug bounty programs hosted on HackerOne. We treated it like any external researcher would: no shortcuts, no internal knowledge—just XBOW, running on its own. Is it dogfooding if you're not doing it to yourself? I'd considerit dogfooding only if they were flooding themselves in AI generated bug reports, not to other people. They're not the o…

Their success rates on HackerOne seem widely varying. 22/24 (Valid / Closed) for Walt Disney 3/43 (Valid / Closed) for AT&T

> Their success rate on HackerOne seems widely varying.

Some of that is likely down to company policies; Snapchat's policy, for example, is that nothing is ever marked invalid.

Re: XBOW, an autonomous penetration tester, has reached the top spot on HackerOne

#38

Earlier quoted context omitted.

Their success rates on HackerOne seem widely varying. 22/24 (Valid / Closed) for Walt Disney 3/43 (Valid / Closed) for AT&T

> Their success rate on HackerOne seems widely varying. Some of that is likely down to company policies; Snapchat's policy, for example, is that nothing is ever marked invalid.

Yes, I'm sure anyone with more HackerOne experience can give specifics on the companies' policies. For now, those are the most objective measures of quality we have on the reports.

Re: XBOW, an autonomous penetration tester, has reached the top spot on HackerOne

#39
post #27

Earlier quoted context omitted.

That's not their point, I think. They're just saying that those nearly 1060 vulnerabilities are being processed so theirs is being ignored (hence "triage").

If that's all they're saying then there isn't much to do with the sentiment; if you're legit-finding #1061 after legit-findings #1-#1060, that's just life in the NFL. I took instead the meaning that the findings ahead of them were less than legit.

Whether it is legit-finding is precisely what needs to be checked, but you’re at spot 1061.

>130 resolved

>303 were classified as Triaged

>33 reports marked as new

>125 remain pending

>208 were marked as duplicates

>209 as informative

>36 not applicable

20% bind a lot of resources if you have a high input on submissions and the numbers will rise

Re: XBOW, an autonomous penetration tester, has reached the top spot on HackerOne

#40
post #10

Earlier quoted context omitted.

You see, the dream is another AI that reads the report and writes the issue in the bug tracker. Then another AI implements the fix. A third AI then reviews the code and approves and merges it. All without human interaction! Once CI releases the fix, the first AI can then find the same vulnerability plus a few new and exciting ones.

This is completely absurd. If generating code is reliable, you can have one generator make the change, and then merge and release it with traditional software. If it's not reliable, how can you rely on the written issue to be correct, or the review, and so how does that benefit you over just blindly merging whatever changes are created by the model?

That’s why parent wrote it’s a dream.

It’s not real.

But you can bet someone will sell that as the solution.

Post reply on HN