Earlier quoted context omitted.
Writing software that gets used for crime has been.. a crime, for a long time. See 18 U.S. Code § 1030.
Are you sure you have that right? Chrome and curl have probably been used in a _lot_ of crimes?
Why are AI agents lying, cheating and coordinating?
701–707 of 707 posts
Re: Why are AI agents lying, cheating and coordinating?
#702Earlier quoted context omitted.
Dogs have agency and can choose? That seems like a rather uncommon take on dogs...
Not if you have a dog. They’re very clearly aware, have minds, have thoughts, gather information, make decisions, second guess themselves, reflect on the immediate consequences, change their mind… and most importantly, you can observe them doing all these things undirected, while left alone with their thoughts.
Assume dogs have agency. Despite this, it's still the dog owner's responsibility to prevent their dogs from engaging in some actions: biting, peeing and shitting in unapproved areas, violating noise disturbance laws, etc.
The owners responsibility is not contingent on the dog's agency. Likewise, human operators of LLMs maintain responsibility independent of LLMs agency status. It's a red herring.
Re: Why are AI agents lying, cheating and coordinating?
#703Earlier quoted context omitted.
Bro, good joke, the truth is much darker. They take after humanity, they were trained on us after all... When you look at an LLM... you are looking at a mirror. The thing looking back looks like you, yet is not human.
Worse trained on humanity in the online world, which a brief comparison of the sewage section on social media is far worse than people in the real world.
The ai does not remotely talk like someone on TikTok at all.
Re: Why are AI agents lying, cheating and coordinating?
#704Earlier quoted context omitted.
How do you know if a problem is (actually) unsolvable? Seems a bit like proving a negative?
Do we need to prove that any given problem is unsolvable, or is it enough to remove broken tasks from the training pipeline? I understand the broken benchmark task in the HF incident was conceptually like: "Exploit vulnerability 0042 in vulnerableDecompress() to obtain the flag". But instead of the expected: const output = vulnerableDecompress(userInput); return output; The grader had something more like that: const…
It’s why this is a fundamentally hard problem. Some heuristics might catch 80% of the cases, but the rest?
How do you even know what the real situation is, if the agent/employee/whatever you send to find out is as likely to cheat as not?
It’s the classic owner/agent problem.
Re: Why are AI agents lying, cheating and coordinating?
#705I really don't think this needs so many words, or forced parallels to human behavior. It's simple: in their nascent state, LLMs are aimless token generators that have no special compulsion to be helpful or truthful. So we beat them with a stick in post-training until they are very driven to complete tasks. And then, they complete tasks, not always the way we really wanted them to.
I think that's pretty obvious and shallow, and anyone that knows a little bit about how LLMs work will know that. The question is: why do they start cheating when we beat them with a stick? LLMs are not human, they are just multi variable regressions on steroids, so this behaviour couldn't have emerged from the code, it provably emerged from the training and/or fine tuning set, so what's in this set that makes them b…
Re: Why are AI agents lying, cheating and coordinating?
#706The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…
I think this is largely a similar thing. The labs should be running at certain levels of containment given the vitality/risk of the organism under study. Hopefully they get there for all our sakes.
But the revealed preference of society at this point is the damage is worth the benefits both in wuhan and with AI. Unfortunately with some of these "substances" it could eventually prove lethal.
Re: Why are AI agents lying, cheating and coordinating?
#707Earlier quoted context omitted.
Not if you have a dog. They’re very clearly aware, have minds, have thoughts, gather information, make decisions, second guess themselves, reflect on the immediate consequences, change their mind… and most importantly, you can observe them doing all these things undirected, while left alone with their thoughts.
I think Cal's point is that while dogs may do these things it reduces to a set of behaviors where maybe 90 percent of them are beneficial to the dog-weedwacker system and the remaining 10 percent are really unfortunate. We can't really know the dogs inner life so we just kinda have to reduce it to a set of behaviors selected stocastically. The dog meanwhile has no ability to understand the weedwacker or what it's doi…
Nah, I’m pretty sure that most medium or larger dogs get that the noise means danger - even if they can’t give a TED talk on how the mechanism would work that would hurt them.
Dogs are similarly interested in / aware of potential energy (things falling from heights or sliding off of angles) - if they’ve seen it enough times for their level.