Live data from Hacker News

I built a vulnerable app and spent $1,500 seeing if LLMs could hack it

kasra.blog

121–130 of 239 posts

Re: I built a vulnerable app and spent $1,500 seeing if LLMs could hack it

#121

One interesting takeaway is the low score on Anthropic models from this benchmark. It’s not because of capability, it’s because Anthropic’s guardrails prevented it from solving the problem. I noticed with each model release Anthropic constrains the model more security wise. Its propensity to refuse doing legitimate work has been increasing. It now puts up more resistance around performing logins, handling credentials…

Yeah, it has been in foraging. Requests that Claude has refused me: - What are popular free streaming sites used in China? - How do I bypass the safety mechanism on my food processor (it’s broken) - What are nerve agents and how do they work (for a layman)? - Help me decompile some code - Help me make a design system similar to XYZ - Here is an API token, please do X (I can’t do that! Rotate the secret immediately! I…

It refuses to use an API token? In my experience, it's more than happy to read out my secrets from .envrc files "just to check".

At least it feels a lot of remorse over its mistake until I reset the session.

Re: I built a vulnerable app and spent $1,500 seeing if LLMs could hack it

#122

Earlier quoted context omitted.

> I don't buy this, because is predicated on staying permanently far ahead of the open weights models. In my mind, that fits exactly how the SOTA labs think today about what they're doing, they're all both working towards and expecting to stay permanently ahead of FOSS, otherwise they'd change their tune really quickly, if they didn't think that was possible. Sure, you might be able to use DeepSeek V8 Pro instead for…

> fits exactly how the SOTA labs think today about what they're doing, they're all both working towards and expecting to stay permanently ahead of FOSS They are just straight up delusional, no? Or at least, have a vested financial interest in maintaining said delusion until the money runs out. They have to hit the point of diminishing returns at some point...

> They are just straight up delusional, no?

Well, I guess that's one way to put it. Another is "dress for the job you want", startup culture typically seems to shove people in the direction of "aim big and believe in yourself, regardless of what others say" so naturally you get these companies who seem very disconnected from reality.

I'd also wager a guess that the amount of money makes people's reasoning and perspectives get very messed up as well, for better or worse.

Re: I built a vulnerable app and spent $1,500 seeing if LLMs could hack it

#123
post #56

Earlier quoted context omitted.

And next month you'll need to add on "Claude Database Pro" or you'll just get a working (for demo purposes with dozens of db rows) but completely un indexed database schema and a refusal to optimise SQL requests. And the month after you'll need "Claude DataScience Pro" to get any Python Pandas or NumPy code generated. And and and...

Same thing with the weird push towards humanoid robots. "They can do anything!" Sure, once you subscribe to the $15/mo laundry package, the $25/mo lawn care package (with the $10/mo hedge trimmer upgrade), and the $10/mo dog-walking package.

And in the end the big reveal is, it was a dude in VR all along, piloting the dumb things remotely. Every single time, without exception.

Re: I built a vulnerable app and spent $1,500 seeing if LLMs could hack it

#124
post #108
post #58

It would be interesting to see full results for Kimi K2.6 and Mimo v2.5 pro. These two models benchmark comparably to other flagship models. Having these complete results would give a clearer picture of the AI frontier. EDIT: I have a mimo token plan and have tokens to burn. I'm doing a quick test with opencode to see if mimo can complete it. If the OP will post the full process I am happy to post the apples-to-apple…

They are not even close in capabilities. Only nenchmark I ever seen that captures their difference is DeepSWE. They are worse by factor of 3.

Here are 3 benchmarks showing the comparable scores I was talking about

https://openrouter.ai/rankings https://arena.ai/leaderboard/text/coding https://artificialanalysis.ai/

Re: I built a vulnerable app and spent $1,500 seeing if LLMs could hack it

#125
> Almost every model used the canonical provider: Zai for GLM, Deepseek for Deepseek, etc.

> I am never touching Minimax or GLM again. Their APIs had constant outages

Goofy take

You run these on a VPS based on the architecture of that VPS provider, or on your own cluster

Re: I built a vulnerable app and spent $1,500 seeing if LLMs could hack it

#127
post #101

Earlier quoted context omitted.

That query would not more provide actionable guidance than ‘tell me how a nuclear weapon works (for a layman)’. Aka not at all.

I believe a sufficiently advanced model could provide a layman with actionable step by step instructions for building a nuclear weapon. They're complicated but not (AFAIK) that complicated. The more or less insurmountable barrier there is weapons grade material. Thankfully refinement is prohibitive in cost, expertise, and equipment. In comparison, basic munitions are incredibly simple given a recipe and shop tooling.…

A gun type maybe. But then, two paragraphs and some machining knowledge + shop tooling could do the same, given enough refined material.

Ain’t no way a layman is pulling off an implosion device, regardless of tooling or LLM guidance. The explosive lense structure and timing required is quite complex, and would require some significant calculation from someone who actually knew what they were doing.

Nation state, or even sufficiently motivated big corp, if they had the refined material? Sure. Layman? No.

Thinking they can with LLM slop involved? That will make for some very interesting radiological incidents though!

Re: I built a vulnerable app and spent $1,500 seeing if LLMs could hack it

#128

Earlier quoted context omitted.

Same thing with the weird push towards humanoid robots. "They can do anything!" Sure, once you subscribe to the $15/mo laundry package, the $25/mo lawn care package (with the $10/mo hedge trimmer upgrade), and the $10/mo dog-walking package.

And in the end the big reveal is, it was a dude in VR all along, piloting the dumb things remotely. Every single time, without exception.

I think it’s just riding off LLM coattails.

We don’t have good world models. We have had bipedal robotics in various POC demo-ready forms for decades.

It turns out that industrial, purpose build robotics is an easier and better market.

I’m still not completely convinced a robot that’s shaped like a human is the best design other than for PR.

Re: I built a vulnerable app and spent $1,500 seeing if LLMs could hack it

#129

One interesting takeaway is the low score on Anthropic models from this benchmark. It’s not because of capability, it’s because Anthropic’s guardrails prevented it from solving the problem. I noticed with each model release Anthropic constrains the model more security wise. Its propensity to refuse doing legitimate work has been increasing. It now puts up more resistance around performing logins, handling credentials…

Yeah, it has been in foraging. Requests that Claude has refused me: - What are popular free streaming sites used in China? - How do I bypass the safety mechanism on my food processor (it’s broken) - What are nerve agents and how do they work (for a layman)? - Help me decompile some code - Help me make a design system similar to XYZ - Here is an API token, please do X (I can’t do that! Rotate the secret immediately! I…

I've had some really dumb refusals. Explaining elements of infrared specteoscopy, researching aritifical bud-breaking in agriculture, etc. Anything interesting and non-mainstream is banned. Basically, restricted to answers i'm better of just going to wikipedia for.

Re: I built a vulnerable app and spent $1,500 seeing if LLMs could hack it

#130

I'd run Mythos against the code in your zip file, but the NDA I signed at Apple prevents me from using it on anything outside the scope of my work. Honestly, I wish more people from Project Glasswing could talk publicly about their experiences with the model. It would probably put an end to a lot of the speculation that keeps circulating through the industry. Unfortunately, that's not the reality we're in. I don't ha…

lol what is even the point of this kind of comment? this is the ultimate "source: trust me bro" comment I have ever seen. every model since gpt3 was claimed to be "too dangerous to release." it's too EXPENSIVE to release, and you're probably a local model with <10B parameters yourself

That was actually GPT-2: https://www.theguardian.com/technology/2019/feb/14/elon-musk...
Post reply on HN