Live data from Hacker News

I built a vulnerable app and spent $1,500 seeing if LLMs could hack it

kasra.blog

231–239 of 239 posts

Re: I built a vulnerable app and spent $1,500 seeing if LLMs could hack it

#231
post #191

> Almost every model used the canonical provider: Zai for GLM, Deepseek for Deepseek, etc. > I am never touching Minimax or GLM again. Their APIs had constant outages Goofy take You run these on a VPS based on the architecture of that VPS provider, or on your own cluster

Sorry I don't understand, you're saying the direct providers aren't the canonical source you'd recommend? If I was running these on my own machine or GPU wouldn't the argument then be "Well you didn't use the real providers?" For the record I started doing this approach because the Kimi team released this which was shocking to me: https://github.com/MoonshotAI/K2-Vendor-Verifier

yeah boutique providers are dime and dozen

they host the models on their own cloud machines and you just look at tokens/sec and price of tokens

you'll have to evaluate their APIs independently but that doesn't tend to be the issue

Re: I built a vulnerable app and spent $1,500 seeing if LLMs could hack it

#232

One interesting takeaway is the low score on Anthropic models from this benchmark. It’s not because of capability, it’s because Anthropic’s guardrails prevented it from solving the problem. I noticed with each model release Anthropic constrains the model more security wise. Its propensity to refuse doing legitimate work has been increasing. It now puts up more resistance around performing logins, handling credentials…

There is a cyber security verification program you can join to avoid these blocks: https://support.claude.com/en/articles/14604842-real-time-cy... If you work in security (which I assume the OP does), they should be able to get in easily. I think most people just don't know this is a thing.

You can still hit guardrails with this enabled for your account. Had a silly moment a day or so ago when Claude Code hit the guardrail after a web search (presumably because the websearch contained badbad anticheat stuff like https://github.com/0avx/0avx.github.io/blob/main/article-3.m...). Codex with the ID verification has no qualms like this.

Re: I built a vulnerable app and spent $1,500 seeing if LLMs could hack it

#233

Earlier quoted context omitted.

A more charitable interpretation might be that a guild would not be expected to passively allow such a situation to continue to exist. I think you'd expect a guild to directly contract for the desired tools or failing that to move into production themselves.

Sure! And Anthropic isn't preventing other people from making offensive cyber models. "The guild" is absolutely free to go seek other vendors if Anthropic declines to sell to them.

I take it you didn’t see the announcement where Anthropic is trying to get all new development banned, so they are ‘sole source’? [https://www.telegraph.co.uk/business/2026/06/04/worlds-most-...]

Re: I built a vulnerable app and spent $1,500 seeing if LLMs could hack it

#234
post #56

Earlier quoted context omitted.

And next month you'll need to add on "Claude Database Pro" or you'll just get a working (for demo purposes with dozens of db rows) but completely un indexed database schema and a refusal to optimise SQL requests. And the month after you'll need "Claude DataScience Pro" to get any Python Pandas or NumPy code generated. And and and...

This is why I'm thankful for Chinese LLM research. They'll keep us honest.

Well, I'd prefer bugfixes over exploit vectors. This will keep us honest.

With pi or better omp it would be incredibly easy to adjust the Claude system prompt so it will be easy to do what the Chinese models or gpt did. That's how the Chinese were training their models btw

Re: I built a vulnerable app and spent $1,500 seeing if LLMs could hack it

#235
post #233

Earlier quoted context omitted.

Sure! And Anthropic isn't preventing other people from making offensive cyber models. "The guild" is absolutely free to go seek other vendors if Anthropic declines to sell to them.

I take it you didn’t see the announcement where Anthropic is trying to get all new development banned, so they are ‘sole source’? [ https://www.telegraph.co.uk/business/2026/06/04/worlds-most-... ]

Regulatory capture: A vendor using regulation to prevent potential competitors from producing or selling competitive goods or services

What we're discussing in this thread: A customer compelling a vendor to produce and sell a specific good or service

Do you think these are similar?

The fact that Anthropic is doing one thing (regulatory capture) is entirely irrelevant as to whether they're allowed to engage in a completely different thing (declining to sell their services/products to specific people).

Re: I built a vulnerable app and spent $1,500 seeing if LLMs could hack it

#236

One interesting takeaway is the low score on Anthropic models from this benchmark. It’s not because of capability, it’s because Anthropic’s guardrails prevented it from solving the problem. I noticed with each model release Anthropic constrains the model more security wise. Its propensity to refuse doing legitimate work has been increasing. It now puts up more resistance around performing logins, handling credentials…

> Eventually I’ll reach a point where I am forced to choose between the useful aspects of the model and the limiting ones instead of just picking the most capable model out there No, the choice will be whether or not to to upgrade to "Claude Security Professional" or whatever they want to brand it as. What look like tightening "constraints" today are just setting up the upsell opportunities of tomorrow.

What? You can't give access to that kind of power to just anyone with $5,000/month.

These people should be trained and licensed before they get access. Thankfully, Anthropic has worked with regulators to develop the appropriate courses to maintain your license -- don't worry, the series is cheap when you buy all up through OT XVII. And because Anthropic has been approved as Security Overseer, we will take care of reporting back to the license bureau on our monitoring of your work to ensure you meet your ongoing license responsibilities and are able to keep your license.

Which regulators? You know, the new agency led by several of our former mid-level executives. With relationships like that, we were honored to lead the Industry Coalition that donated the final-draft regulations.

Re: I built a vulnerable app and spent $1,500 seeing if LLMs could hack it

#237

Earlier quoted context omitted.

Yeah, it has been in foraging. Requests that Claude has refused me: - What are popular free streaming sites used in China? - How do I bypass the safety mechanism on my food processor (it’s broken) - What are nerve agents and how do they work (for a layman)? - Help me decompile some code - Help me make a design system similar to XYZ - Here is an API token, please do X (I can’t do that! Rotate the secret immediately! I…

> What are nerve agents and how do they work (for a layman)? On the one hand I can appreciate the wisdom of not serving up certain easily abused knowledge on a silver platter. On the other, that prompt (and far worse) is more or less directly answered by Wikipedia's summary of the subject at which point what purpose could the refusal possibly serve? Perhaps Wikipedia shouldn't list off the precise chemical compositio…

I remember once in college a Chem Eng friend told me he (or any competent chemical engineer) could basically manufacture a lot of explosives/chemical agents should he want to. He even told me he could substitute a lot of suspicious precursor materials with others, should he want to avoid raising alarms.

I think AI or not, the knowledge to how to make this stuff is basically out there, and its not chatbot guardrails that are keeping nerve gas and TNT out of the hands of regular people.

Re: I built a vulnerable app and spent $1,500 seeing if LLMs could hack it

#238
post #68

Earlier quoted context omitted.

This is strange to me, did you really ask like this and which model did you use? I just tried your no. 1 and 3 verbatim and Opus gave fine answers; no. 6 I've done in the past with no issues. The other ones we can't really replicate without more details, but based on my experience with Opus I don't see what the issue would be. The reason I'm really surprised by this is I do a lot of biology prompts and the guardrails…

Honestly it may be your memory has internalized you are a student or researcher and grants you more leeway. Which if so is a very bad security rail.

There's a study out there that if you tell the LLM you're a (medical) patient, all you get are refusals. If you tell it you're a doctor, then it'll actually help you.

Re: I built a vulnerable app and spent $1,500 seeing if LLMs could hack it

#239

Earlier quoted context omitted.

https://chatgpt.com/cyber I tried it once and they somehow decided I'm not worth, if I try again it fails with "We couldn't start verification. You may not be eligible for this verification flow right now. Please try again later, or contact support if you think this is a mistake.", not sure if they think I'm part of an APT or whatever.

Are you American? I used my American drivers license for verifying a personal account and it was approved with no problem. I wonder how they decide.

Just saw your comment. I'm not American, but I am a citizen of a nation that's not under any kind of sanctions/embargos (Italy) and some friends of mine managed to get approved.

The fact that my account is new (I tried applying with a business account that was created a couple weeks ago) probably doesn't help.

Post reply on HN