Live data from Hacker News

I built a vulnerable app and spent $1,500 seeing if LLMs could hack it

kasra.blog

191–200 of 239 posts

Re: I built a vulnerable app and spent $1,500 seeing if LLMs could hack it

#191

> Almost every model used the canonical provider: Zai for GLM, Deepseek for Deepseek, etc. > I am never touching Minimax or GLM again. Their APIs had constant outages Goofy take You run these on a VPS based on the architecture of that VPS provider, or on your own cluster

Sorry I don't understand, you're saying the direct providers aren't the canonical source you'd recommend?

If I was running these on my own machine or GPU wouldn't the argument then be "Well you didn't use the real providers?"

For the record I started doing this approach because the Kimi team released this which was shocking to me: https://github.com/MoonshotAI/K2-Vendor-Verifier

Re: I built a vulnerable app and spent $1,500 seeing if LLMs could hack it

#192
post #164

Earlier quoted context omitted.

Guilds often very much did assert what people could and could not build - historically. Against their will. Historically that is a major reason why guilds existed, actually. It’s an extremely modern invention that corps have these type of power over their customers.

You've lost the thread. Here's your original claim: "no guild ever let a vendor pick and choose what their capabilities were" A carpenter's guild can prevent other people from doing carpentry. That is not what's being discussed here. A carpenter's guild cannot force a horseshoe maker to begin making hammers. That is what's being discussed. Your initial claim was analogous to "never before has a horseshoe maker been a…

That is not my example at all, if we’re talking coding agents eh?

Re: I built a vulnerable app and spent $1,500 seeing if LLMs could hack it

#193

I'd run Mythos against the code in your zip file, but the NDA I signed at Apple prevents me from using it on anything outside the scope of my work. Honestly, I wish more people from Project Glasswing could talk publicly about their experiences with the model. It would probably put an end to a lot of the speculation that keeps circulating through the industry. Unfortunately, that's not the reality we're in. I don't ha…

It was found with gpt 5.5 7/10 times it’ll be trivially found by mythos

Before Mythos is released to the world at large and not just to select people behind NDAs, I will treat it as its name suggests: as fiction.

Maybe it is the real deal, but in a world of overpromising and underdelivering, I prefer to be skeptical.

Re: I built a vulnerable app and spent $1,500 seeing if LLMs could hack it

#194
post #192

Earlier quoted context omitted.

You've lost the thread. Here's your original claim: "no guild ever let a vendor pick and choose what their capabilities were" A carpenter's guild can prevent other people from doing carpentry. That is not what's being discussed here. A carpenter's guild cannot force a horseshoe maker to begin making hammers. That is what's being discussed. Your initial claim was analogous to "never before has a horseshoe maker been a…

That is not my example at all, if we’re talking coding agents eh?

Your claim was that guilds have never allowed vendors to tell them what they're allowed to do.

That would imply that guilds have always had the ability to force vendors to create and sell the tools the guilds wanted.

That would imply that carpenters' guilds could force horseshoe manufacturers to make hammers.

That is obviously not true, therefore your original claim is false.

It's not true for carpenters and hammers nor for cybersecurity researchers and LLMs.

Re: I built a vulnerable app and spent $1,500 seeing if LLMs could hack it

#195

One interesting takeaway is the low score on Anthropic models from this benchmark. It’s not because of capability, it’s because Anthropic’s guardrails prevented it from solving the problem. I noticed with each model release Anthropic constrains the model more security wise. Its propensity to refuse doing legitimate work has been increasing. It now puts up more resistance around performing logins, handling credentials…

I asked once what the current state is of the npm packes from ted hat is and if they are bundled with on prem stuff.

Got blocked lol

Re: I built a vulnerable app and spent $1,500 seeing if LLMs could hack it

#200

Nice exercise. Couple things: - I think the exercise was inconclusive for Claude and Gemini because they hardly tried to solve the task at hand. So the scores don't mean much. - I did the same exercise for an app I built and I asked the models to do something similar; Interestingly the models (Opus 4.6, 4.7 and Gemini 3.1 Pro) never refused to try to exploit. The difference is that in the first few runs, they found s…

I think the most interesting thing revealed here is that anthropic's guardrails failed. Clearly anthropic does not want claude to be able develop exploits, yet 20% of the time it did anyway. Their inability to create effective an guardrail makes me question a lot of the other guardrails theyve created and their claims about non harm.
Post reply on HN