Earlier quoted context omitted.
If that's all they're saying then there isn't much to do with the sentiment; if you're legit-finding #1061 after legit-findings #1-#1060, that's just life in the NFL. I took instead the meaning that the findings ahead of them were less than legit.
Whether it is legit-finding is precisely what needs to be checked, but you’re at spot 1061. >130 resolved >303 were classified as Triaged >33 reports marked as new >125 remain pending >208 were marked as duplicates >209 as informative >36 not applicable 20% bind a lot of resources if you have a high input on submissions and the numbers will rise
XBOW, an autonomous penetration tester, has reached the top spot on HackerOne
41–50 of 128 posts
Re: XBOW, an autonomous penetration tester, has reached the top spot on HackerOne
#42Earlier quoted context omitted.
Moreover, I don't think XBOW is likely generating the kind of slop beg bounty people generate. There's some serious work behind this.
Do you have sources for if we want to learn more?
Some of my favorites from what we've released so far:
- Exploitation of an n-day RCE in Jenkins, where the agent managed to figure out the challenge environment was broken and used the RCE exploit to debug the server environment and work around the problem to solve the challenge: https://xbow.com/#debugging--testing--and-refining-a-jenkins...
- Authentication bypass in Scoold that allowed reading the server config (including API keys) and arbitrary file read: https://xbow.com/blog/xbow-scoold-vuln/
- The first post about our HackerOne findings, an XSS in Palo Alto Networks GlobalProtect VPN portal used by a bunch of companies: https://xbow.com/blog/xbow-globalprotect-xss/
Re: XBOW, an autonomous penetration tester, has reached the top spot on HackerOne
#43"XBOW is an enterprise solution. If your company would like a demo, email us at info@xbow.com." Like any "AI" article, this is an ad. If you are willing to tolerate a high false positive rate, you can as well use Rational Purify or various analyzers.
You should come to my upcoming BlackHat talk on how we did this while avoiding false positives :D https://www.blackhat.com/us-25/briefings/schedule/#ai-agents...
I know you've been on HN for awhile, and that you're doing interesting stuff; HN just has a really intense immune system against vendor-y stuff.
Re: XBOW, an autonomous penetration tester, has reached the top spot on HackerOne
#44Receiving hundreds of AI generated bug reports would be so demoralizing and probably turn me off from maintaining an open source project forever. I think developers are going to eventually need tools to filter out slop. If you didn’t take the time to write it, why should I take the time to read it?
Re: XBOW, an autonomous penetration tester, has reached the top spot on HackerOne
#45Earlier quoted context omitted.
> Their success rate on HackerOne seems widely varying. Some of that is likely down to company policies; Snapchat's policy, for example, is that nothing is ever marked invalid.
Yes, I'm sure anyone with more HackerOne experience can give specifics on the companies' policies. For now, those are the most objective measures of quality we have on the reports.
Re: XBOW, an autonomous penetration tester, has reached the top spot on HackerOne
#46Earlier quoted context omitted.
You should come to my upcoming BlackHat talk on how we did this while avoiding false positives :D https://www.blackhat.com/us-25/briefings/schedule/#ai-agents...
You should publish the paper quietly here (I'm a Black Hat reviewer, FWIW) so people can see where you're coming from. I know you've been on HN for awhile, and that you're doing interesting stuff; HN just has a really intense immune system against vendor-y stuff.
I'll see if I can get time to do a paper to accompany the BH talk. And hopefully the agent traces of individual vulns will also help.
Re: XBOW, an autonomous penetration tester, has reached the top spot on HackerOne
#47Earlier quoted context omitted.
You should publish the paper quietly here (I'm a Black Hat reviewer, FWIW) so people can see where you're coming from. I know you've been on HN for awhile, and that you're doing interesting stuff; HN just has a really intense immune system against vendor-y stuff.
Yeah, it's been very strange being on the other side of that after 10 years in academia! But it's totally reasonable for people to be skeptical when there's a bunch of money sloshing around. I'll see if I can get time to do a paper to accompany the BH talk. And hopefully the agent traces of individual vulns will also help.
Re: XBOW, an autonomous penetration tester, has reached the top spot on HackerOne
#48Earlier quoted context omitted.
Yeah, it's been very strange being on the other side of that after 10 years in academia! But it's totally reasonable for people to be skeptical when there's a bunch of money sloshing around. I'll see if I can get time to do a paper to accompany the BH talk. And hopefully the agent traces of individual vulns will also help.
J'accuse! You were required to do a paper for BH anyways! :)
> White Paper/Slide Deck/Supporting Materials (optional)
> • If you have a completed white paper or draft, slide deck, or other supporting materials, you can optionally provide a link for review by the board.
> • Please note: Submission must be self-contained for evaluation, supporting materials are optional.
> • PDF or online viewable links are preferred, where no authentication/log-in is required.
(From the link on the BHUSA CFP page, which confusingly goes to the BH Asia doc: https://i.blackhat.com/Asia-25/BlackHat-Asia-2025-CFP-Prepar... )
Re: XBOW, an autonomous penetration tester, has reached the top spot on HackerOne
#49Receiving hundreds of AI generated bug reports would be so demoralizing and probably turn me off from maintaining an open source project forever. I think developers are going to eventually need tools to filter out slop. If you didn’t take the time to write it, why should I take the time to read it?
Re: XBOW, an autonomous penetration tester, has reached the top spot on HackerOne
#50Earlier quoted context omitted.
That's not their point, I think. They're just saying that those nearly 1060 vulnerabilities are being processed so theirs is being ignored (hence "triage").
If that's all they're saying then there isn't much to do with the sentiment; if you're legit-finding #1061 after legit-findings #1-#1060, that's just life in the NFL. I took instead the meaning that the findings ahead of them were less than legit.
I see what you're saying but I think a more charitable interpretation can be made. They may be amazed that so many bug reports are being generated by such a reputable group. Looking at your initial reply, perhaps a more constructive comment could be one that joins them in excitement (even if that assumption is erroneous) and expanding on why you think it is exciting (e.g. this group's reputation for quality).