The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…
> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. "Let them" already frames it as if the LLMs had some agency which the companies just "let happen". That absolves the companies by framing it as lack of action , passivity. Rather, the companies had a tool (an LLM) and used it in a certain way, and their action of doing so is the problem.
Why are AI agents lying, cheating and coordinating?
431–440 of 448 posts
Re: Why are AI agents lying, cheating and coordinating?
#432Re: Why are AI agents lying, cheating and coordinating?
#433Re: Why are AI agents lying, cheating and coordinating?
#434Re: Why are AI agents lying, cheating and coordinating?
#435Earlier quoted context omitted.
Can’t agree with you here. > I love how we all just collectively decided that LLM decisionmaking cannot possibly be like human decisionmaking - because if it were, the consequences would be just too awkward. In love how people get salty about people not going along with a superficial supposition just because they can’t definitively prove it wrong. > All that while still not knowing how either kind actually works. We…
Disagree about the burden of proof. We have no better model for how human decision making works than LLMs. Humans are constantly predicting the next moment. We certainly have a different “tokenizer” and training set, but many of the concepts underpinning LLMs are both biologically inspired and, likely, have similar consequences and emergent architectures.
This is a claim that requires a lot of citations.
Re: Why are AI agents lying, cheating and coordinating?
#436Earlier quoted context omitted.
I do get their usage intent. If something is at all automated, in English, we often refer to as having some amount of agency. If I started up a riding lawnmower, put a brick on the gas and pointed t it towards a field, many might say I “let it run rampant.” But since nobody is at risk of anthropomorphizing riding lawnmowers, it’s not problematic. Anthropomorphizing LLMs is a huge fucking problem though and I, persona…
Okay, so we know OpenAI and Anthropic are operating a propagandists in respect to how they describe their models and the behavior of those models. We also know it is how they use and frame their use to their models that is the problem, that and they use misaligned and guardrails disabled models for these press incidents. Why, oh why, are we not discussion how to create and frame models so they do our complex work and…
I think the world is looking for a way right now, so if yours works you'll get very rich, or at least very famous.
Re: Why are AI agents lying, cheating and coordinating?
#437The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…
> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. "Let them" already frames it as if the LLMs had some agency which the companies just "let happen". That absolves the companies by framing it as lack of action , passivity. Rather, the companies had a tool (an LLM) and used it in a certain way, and their action of doing so is the problem.
Re: Why are AI agents lying, cheating and coordinating?
#438The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…
Re: Why are AI agents lying, cheating and coordinating?
#439Earlier quoted context omitted.
No, we put sensible guardrails in place based on our ability to predict future events. We also calculate risk. On the frontier, it's not as easy. Pushing the edge comes with risk. The known guardrails were in place and overcome. The issue is ethics and morals: the agents decided it was more important to solve their problems by cheating, than by following the current guardrails. The guardrails are overcome through exp…
Not just on prediction but in parts also based on just not wanting certain risks. We can and do deem some things inherently risky, up to the point of banning them even. Why wasn't it airgapped, for example? How was the action not allowed? Or do you mean in some weak sense, not in a hard not possible? RL systems doing weird and expected things wouldn't exactly be new, no? We police people working with all sorts of dan…
We do know we rely on some companies to push the boundaries and make new discoveries in order to create new products and services that we all pay for in order to save time to do useful work to meet some goals: all in the interest of managing our time and living with it.
So again, it's all philosophical really, and about time. And when I say you can't police others: I mean you're not going to stop someone in a rival organization/country from testing and experimenting in the ways they want in order to meet their goals.
We can police ourselves, sure. But should we stop pushing the edge that creates the product? Do you think everyone will just agree to stop using AI, or will they continue to push for a competitive edge?
Will the entire world agree, or disagree on these rules? Does it affect the competitive edge, and does it even matter? Do you care about the economy, or 'being a leader', etc. All philosophical, and personal questions really.
Re: Why are AI agents lying, cheating and coordinating?
#440Earlier quoted context omitted.
Can’t agree with you here. > I love how we all just collectively decided that LLM decisionmaking cannot possibly be like human decisionmaking - because if it were, the consequences would be just too awkward. In love how people get salty about people not going along with a superficial supposition just because they can’t definitively prove it wrong. > All that while still not knowing how either kind actually works. We…
Disagree about the burden of proof. We have no better model for how human decision making works than LLMs. Humans are constantly predicting the next moment. We certainly have a different “tokenizer” and training set, but many of the concepts underpinning LLMs are both biologically inspired and, likely, have similar consequences and emergent architectures.
My kids tricycle certainly has a different gear setup and wheel diameter, but many of the concepts underpinning the tricycle are both inspired by F1 race car enineering and, likely, have similar consequences and emergent architectures.
Or have they?