Live data from Hacker News

XBOW, an autonomous penetration tester, has reached the top spot on HackerOne

xbow.com

41–50 of 128 posts

Re: XBOW, an autonomous penetration tester, has reached the top spot on HackerOne

#41
post #39
post #27

Earlier quoted context omitted.

If that's all they're saying then there isn't much to do with the sentiment; if you're legit-finding #1061 after legit-findings #1-#1060, that's just life in the NFL. I took instead the meaning that the findings ahead of them were less than legit.

Whether it is legit-finding is precisely what needs to be checked, but you’re at spot 1061. >130 resolved >303 were classified as Triaged >33 reports marked as new >125 remain pending >208 were marked as duplicates >209 as informative >36 not applicable 20% bind a lot of resources if you have a high input on submissions and the numbers will rise

I think some context I probably don't share with the rest of this thread is that the average quality of a Hacker One submission is incredibly low. Like however bad you think the median bounty submission is, it's worse; think "people threatening to take you to court for not paying them for their report that they can 'XSS' you with the Chrome developer console".

Re: XBOW, an autonomous penetration tester, has reached the top spot on HackerOne

#42
post #20

Earlier quoted context omitted.

Moreover, I don't think XBOW is likely generating the kind of slop beg bounty people generate. There's some serious work behind this.

Do you have sources for if we want to learn more?

We've got a bunch of agent traces on the front page of the web site right now. We also have done writeups on individual vulnerabilities found by the system, mostly in open source right now (we did some fun scans of OSS projects found on Docker Hub). We have a bunch more coming up about the vulns found in bug bounty targets. The latter are bottlenecked by getting approval from the companies affected, unfortunately.

Some of my favorites from what we've released so far:

- Exploitation of an n-day RCE in Jenkins, where the agent managed to figure out the challenge environment was broken and used the RCE exploit to debug the server environment and work around the problem to solve the challenge: https://xbow.com/#debugging--testing--and-refining-a-jenkins...

- Authentication bypass in Scoold that allowed reading the server config (including API keys) and arbitrary file read: https://xbow.com/blog/xbow-scoold-vuln/

- The first post about our HackerOne findings, an XSS in Palo Alto Networks GlobalProtect VPN portal used by a bunch of companies: https://xbow.com/blog/xbow-globalprotect-xss/

Re: XBOW, an autonomous penetration tester, has reached the top spot on HackerOne

#43
post #34

"XBOW is an enterprise solution. If your company would like a demo, email us at info@xbow.com." Like any "AI" article, this is an ad. If you are willing to tolerate a high false positive rate, you can as well use Rational Purify or various analyzers.

You should come to my upcoming BlackHat talk on how we did this while avoiding false positives :D https://www.blackhat.com/us-25/briefings/schedule/#ai-agents...

You should publish the paper quietly here (I'm a Black Hat reviewer, FWIW) so people can see where you're coming from.

I know you've been on HN for awhile, and that you're doing interesting stuff; HN just has a really intense immune system against vendor-y stuff.

Re: XBOW, an autonomous penetration tester, has reached the top spot on HackerOne

#44

Receiving hundreds of AI generated bug reports would be so demoralizing and probably turn me off from maintaining an open source project forever. I think developers are going to eventually need tools to filter out slop. If you didn’t take the time to write it, why should I take the time to read it?

These aren't like Github Issues reports; they're bug bounty programs, specifically stood up to soak up incoming reports from anonymous strangers looking to make money on their submissions, with the premise being that enough of those reports will drive specific security goals (the scope of each program is, for smart vendors, tailored to engineering goals they have internally) to make it worthwhile.

Re: XBOW, an autonomous penetration tester, has reached the top spot on HackerOne

#45

Earlier quoted context omitted.

> Their success rate on HackerOne seems widely varying. Some of that is likely down to company policies; Snapchat's policy, for example, is that nothing is ever marked invalid.

Yes, I'm sure anyone with more HackerOne experience can give specifics on the companies' policies. For now, those are the most objective measures of quality we have on the reports.

This is discussed in the post – many came down to individual programs' policies e.g. not accepting the vulnerability if it was in a 3rd party product they used (but still hosted by them), duplicates (another researcher reported the same vuln at the same time; not really any way to avoid this), or not accepting some classes of vuln like cache poisoning.

Re: XBOW, an autonomous penetration tester, has reached the top spot on HackerOne

#46
post #43
post #34

Earlier quoted context omitted.

You should come to my upcoming BlackHat talk on how we did this while avoiding false positives :D https://www.blackhat.com/us-25/briefings/schedule/#ai-agents...

You should publish the paper quietly here (I'm a Black Hat reviewer, FWIW) so people can see where you're coming from. I know you've been on HN for awhile, and that you're doing interesting stuff; HN just has a really intense immune system against vendor-y stuff.

Yeah, it's been very strange being on the other side of that after 10 years in academia! But it's totally reasonable for people to be skeptical when there's a bunch of money sloshing around.

I'll see if I can get time to do a paper to accompany the BH talk. And hopefully the agent traces of individual vulns will also help.

Re: XBOW, an autonomous penetration tester, has reached the top spot on HackerOne

#47
post #46
post #43

Earlier quoted context omitted.

You should publish the paper quietly here (I'm a Black Hat reviewer, FWIW) so people can see where you're coming from. I know you've been on HN for awhile, and that you're doing interesting stuff; HN just has a really intense immune system against vendor-y stuff.

Yeah, it's been very strange being on the other side of that after 10 years in academia! But it's totally reasonable for people to be skeptical when there's a bunch of money sloshing around. I'll see if I can get time to do a paper to accompany the BH talk. And hopefully the agent traces of individual vulns will also help.

J'accuse! You were required to do a paper for BH anyways! :)

Re: XBOW, an autonomous penetration tester, has reached the top spot on HackerOne

#48
post #47
post #46

Earlier quoted context omitted.

Yeah, it's been very strange being on the other side of that after 10 years in academia! But it's totally reasonable for people to be skeptical when there's a bunch of money sloshing around. I'll see if I can get time to do a paper to accompany the BH talk. And hopefully the agent traces of individual vulns will also help.

J'accuse! You were required to do a paper for BH anyways! :)

Wait a sec, I thought they were optional?

> White Paper/Slide Deck/Supporting Materials (optional)

> • If you have a completed white paper or draft, slide deck, or other supporting materials, you can optionally provide a link for review by the board.

> • Please note: Submission must be self-contained for evaluation, supporting materials are optional.

> • PDF or online viewable links are preferred, where no authentication/log-in is required.

(From the link on the BHUSA CFP page, which confusingly goes to the BH Asia doc: https://i.blackhat.com/Asia-25/BlackHat-Asia-2025-CFP-Prepar... )

Re: XBOW, an autonomous penetration tester, has reached the top spot on HackerOne

#49

Receiving hundreds of AI generated bug reports would be so demoralizing and probably turn me off from maintaining an open source project forever. I think developers are going to eventually need tools to filter out slop. If you didn’t take the time to write it, why should I take the time to read it?

All of these reports came with executable proof of the vulnerabilities – otherwise, as you say, you get flooded with hallucinated junk like the poor curl dev. This is one of the things that makes offensive security an actually good use case for AI – exploits serve as hard evidence that the LLM can't fake.

Re: XBOW, an autonomous penetration tester, has reached the top spot on HackerOne

#50
post #27

Earlier quoted context omitted.

That's not their point, I think. They're just saying that those nearly 1060 vulnerabilities are being processed so theirs is being ignored (hence "triage").

If that's all they're saying then there isn't much to do with the sentiment; if you're legit-finding #1061 after legit-findings #1-#1060, that's just life in the NFL. I took instead the meaning that the findings ahead of them were less than legit.

> there isn't much to do with the sentiment

I see what you're saying but I think a more charitable interpretation can be made. They may be amazed that so many bug reports are being generated by such a reputable group. Looking at your initial reply, perhaps a more constructive comment could be one that joins them in excitement (even if that assumption is erroneous) and expanding on why you think it is exciting (e.g. this group's reputation for quality).

Post reply on HN