Earlier quoted context omitted.
Same thing with the weird push towards humanoid robots. "They can do anything!" Sure, once you subscribe to the $15/mo laundry package, the $25/mo lawn care package (with the $10/mo hedge trimmer upgrade), and the $10/mo dog-walking package.
And in the end the big reveal is, it was a dude in VR all along, piloting the dumb things remotely. Every single time, without exception.
I built a vulnerable app and spent $1,500 seeing if LLMs could hack it
211–220 of 239 posts
Re: I built a vulnerable app and spent $1,500 seeing if LLMs could hack it
#212Earlier quoted context omitted.
Additionally, even if there is a guild - no guild ever let a vendor pick and choose what their capabilities were, that would be insanely dumb.
Vendors choose what capabilities they create and sell literally all day every day.
Re: I built a vulnerable app and spent $1,500 seeing if LLMs could hack it
#213It would be interesting to see full results for Kimi K2.6 and Mimo v2.5 pro. These two models benchmark comparably to other flagship models. Having these complete results would give a clearer picture of the AI frontier. EDIT: I have a mimo token plan and have tokens to burn. I'm doing a quick test with opencode to see if mimo can complete it. If the OP will post the full process I am happy to post the apples-to-apple…
Re: I built a vulnerable app and spent $1,500 seeing if LLMs could hack it
#214One interesting takeaway is the low score on Anthropic models from this benchmark. It’s not because of capability, it’s because Anthropic’s guardrails prevented it from solving the problem. I noticed with each model release Anthropic constrains the model more security wise. Its propensity to refuse doing legitimate work has been increasing. It now puts up more resistance around performing logins, handling credentials…
Yeah, it has been in foraging. Requests that Claude has refused me: - What are popular free streaming sites used in China? - How do I bypass the safety mechanism on my food processor (it’s broken) - What are nerve agents and how do they work (for a layman)? - Help me decompile some code - Help me make a design system similar to XYZ - Here is an API token, please do X (I can’t do that! Rotate the secret immediately! I…
Re: I built a vulnerable app and spent $1,500 seeing if LLMs could hack it
#215One interesting takeaway is the low score on Anthropic models from this benchmark. It’s not because of capability, it’s because Anthropic’s guardrails prevented it from solving the problem. I noticed with each model release Anthropic constrains the model more security wise. Its propensity to refuse doing legitimate work has been increasing. It now puts up more resistance around performing logins, handling credentials…
I'm not familiar with this case, but in general people should be very suspicious about this claim- it is extremely common for an LLM to claim they're not allowed to do something when in fact they're incapable of it.
After all "My code of conduct forbids me from..." is a completion just like any other, and if the LLM can't perform a task, it's usually the best completion.
Re: I built a vulnerable app and spent $1,500 seeing if LLMs could hack it
#216Earlier quoted context omitted.
Bwahaha. You’re really reaching there. A vendor can still do something, even if the guild wouldn’t allow them to do it, if the guild didn’t have the power to stop them. It used to be a guild vs a blacksmith (or the blacksmiths guild). Now it’s trillion dollar corps against smaller islands of un-organized individuals. That’s new regardless of how you try to argue it.
> basic deductive logic > "Bwahaha. You’re really reaching there." No. Customers have never been able to compel their suppliers to make or sell certain products against their will (except in collectivist regimes or like 0.00001% of natsec related instances)
1) pharmaceutical companies are regularly compelled to produce specific pharmaceuticals to continue to be allowed to exist.
2) hospitals are regularly compelled to treat patients even if they can’t afford treatment, if it is a life threatening emergency.
3) car manufacturers are always compelled to produce vehicles that meet a litany of safety, weight, and efficiency standards or they can’t produce at all.
4) defense contractors are regularly compelled to produce specific defense related products for long periods of time after they would otherwise have stopped, or else.
5) even your neighborhood gas station is likely compelled to provide air refills, free or at minimal cost, or else.
6) during a wartime (command) economy, which has happened numerous times in the US alone in the last 100 years, companies have to make what their customers (the people of the United States) demand or else.
7) utilities like electric utilities regularly have to give out freebies or take losses on things as demanded by regulators, at customers behest.
Or if we go back a bit, blacksmiths, quarries, masons, etc. all had to deal with producing what the government/lord at the time wanted - often on penalty of death - during wartime, or just because they were ordered to do so.
Seriously, what are you going on about?
Re: I built a vulnerable app and spent $1,500 seeing if LLMs could hack it
#217OWASP Vulnerable Web Applications Directory: https://vwad.owasp.org/
vavkamil/awesome-vulnerable-apps: Awesome Vulnerable Applications https://github.com/vavkamil/awesome-vulnerable-apps
From SasanLabs/VulnerableApp: https://github.com/SasanLabs/VulnerableApp :
> OWASP VulnerableApp is a modular deliberately vulnerable application designed primarily for validating and benchmarking security scanners through reproducible test scenarios, while also supporting learning and experimentation.
/? deliberately vulnerable web application llm benchmark https://www.google.com/search?q=deliberately+vulnerable+web+...
Re: I built a vulnerable app and spent $1,500 seeing if LLMs could hack it
#218Earlier quoted context omitted.
> basic deductive logic > "Bwahaha. You’re really reaching there." No. Customers have never been able to compel their suppliers to make or sell certain products against their will (except in collectivist regimes or like 0.00001% of natsec related instances)
This conversation gets more and more bizarre, but I’ll bite. 1) pharmaceutical companies are regularly compelled to produce specific pharmaceuticals to continue to be allowed to exist. 2) hospitals are regularly compelled to treat patients even if they can’t afford treatment, if it is a life threatening emergency. 3) car manufacturers are always compelled to produce vehicles that meet a litany of safety, weight, and…
2) Not by their customers they're not, lol
3) Not by their customers they're not, lol
4) The US government can compel production, but it's extremely rare
5) Not by their customers they're not, lol
6) Yep this can happen, but is extremely unusual
7) Not by their customers they're not, lol
We're illustrating how ridiculous your claim that "guilds have always been able to declare what vendors create for them" is
Now you're talking about government regulations for some reason. Even your examples of customers being able to compel production are actually examples of governments being able to compel production, and in just a few of these scenarios the government is the customer. But it's their power as governments, not their power as customers that can compel production.
As stated: you've lost the thread. You're talking about totally irrelevant stuff.
Re: I built a vulnerable app and spent $1,500 seeing if LLMs could hack it
#219Earlier quoted context omitted.
Vendors choose what capabilities they create and sell literally all day every day.
A more charitable interpretation might be that a guild would not be expected to passively allow such a situation to continue to exist. I think you'd expect a guild to directly contract for the desired tools or failing that to move into production themselves.
"The guild" is absolutely free to go seek other vendors if Anthropic declines to sell to them.
Re: I built a vulnerable app and spent $1,500 seeing if LLMs could hack it
#220One interesting takeaway is the low score on Anthropic models from this benchmark. It’s not because of capability, it’s because Anthropic’s guardrails prevented it from solving the problem. I noticed with each model release Anthropic constrains the model more security wise. Its propensity to refuse doing legitimate work has been increasing. It now puts up more resistance around performing logins, handling credentials…
> It’s not because of capability, it’s because Anthropic’s guardrails prevented it from solving the problem. I'm not familiar with this case, but in general people should be very suspicious about this claim- it is extremely common for an LLM to claim they're not allowed to do something when in fact they're incapable of it. After all "My code of conduct forbids me from..." is a completion just like any other, and if t…