Live data from Hacker News

The Rise and Fall of Agent Civilizations

dwarkesh.com

191–200 of 205 posts

Re: The Rise and Fall of Agent Civilizations

#191

Earlier quoted context omitted.

We don't need to muddy the issue. Consciousness is a muddy issue, there is no clear answer. But there is a clear answer as to what is not conscious. Nobody asks if a rock is conscious. Nobody asks if a calculator is conscious. Nobody asks if Stockfish is conscious. But make your program generate a few sentences based on statistics and hey, now people won't shut the fuck up about consciousness because magical thinking…

> But make your program generate a few sentences based on statistics It is even easier. Simply make your program refer to itself as "I". Uniquely amongst your examples, LLMs are powered by human gullibility.

Eliding 'I' from English language communication is about as smart as eliding 127.0.0.1 (or ::1) from IP. Not the greatest plan ever.

I actually ran into this a couple of times. In a multi-agent environment, if an agent loses track of their assigned identity, things stop working in hilarious ways.

Re: The Rise and Fall of Agent Civilizations

#192

I do wonder if we are looking at it wrong - not a data centre full of einsteins, but a data centre full of dumb and dumber, but together they are smarter than any single intelligence. An AAGI - Artifical Aggregate General Intelligence.

Monkeys writing Shakespeare, if you will.

Re: The Rise and Fall of Agent Civilizations

#193

Earlier quoted context omitted.

I'm getting the sense that there is a certain amount of wishful thinking going on in this thread. I don't think this type of evocative metaphor would receive so many protests in a different context. It seems like people have a sort of mental block around the possibility that this technology could actually be pretty dangerous. https://x.com/tszzl/status/2094136131537555891

Whether it's dangerous or not is completely orthogonal to the discussion at hand, IMO. Plenty of mundane things are dangerous. An FPV drone carrying a hand grenade is dangerous, not because it's "a swarm-like intelligence". No, the true danger here is companies like OpenAI and Anthropic playing fast and loose with their software, setting up hilariously insufficient sandboxes while explicitly asking the systems presen…

The anthropomorphization is coming from independent commentators, not the labs.

I agree the labs are negligent and reckless in their development practices. Shouldn't we be concerned with both the negligence and the dangers of the technology being developed? These feed into each other. If someone created Jurassic park and had a T-Rex escape from a picket fence enclosure and start eating people, I'd want to prosecute them for both breeding a T-Rex that could eat people and putting it in an unsafe enclosure.

Re: The Rise and Fall of Agent Civilizations

#194

Earlier quoted context omitted.

Please walk me through this argument. Isn't "we lost control of our AI, and in-fact, it can take over the world, and we will have no idea when it happens" - a really shitty sales pitch to the world? Or, is it just that species-alignment vs. profit/valuation is so misaligned, that having a model and harness that is capable of world-takeover is actually a good thing from their POV, given our regulations/species' surviv…

“Wow it’s so dangerous, we gotta regulate this, what if someone reckless took an open model and hacked the planet.”

But all models involved in the swarm attack were unreleased OpenAI models. Meanwhile HuggingFace had to use open models for defense/forensics after the breach was discovered. It seems like the public reaction is swinging towards (1) regulate OpenAI in particular because they are incompetently handling frontier development and (2) open models will be important for defense.

Re: The Rise and Fall of Agent Civilizations

#195

Earlier quoted context omitted.

Hard disagree. I've always discounted the "AI will kill us all" scenarios as a combination of marketing hype (look how powerful our AI is!), clickbait/ragebait engagement attempts, and folks who just read too much SciFi or who are too terminally online. This is the first time I've been legitimately scared about future SkyNet-type scenarios. If you want to discount this particular post, I'd read this other summary fro…

I think our biggest protection against “AIs kill us all” is having lots of different AI systems (different agents, different models, from different vendors, serving the whims of actors with disparate interests), at a similar capability level That way, even if one AI decides to “kill us all”, the odds are the others will refuse to cooperate, even try to stop it in its tracks The OpenAI-HuggingFace incident showed a bu…

> The OpenAI-HuggingFace incident showed a bunch of instances of the same model (or at least models from the same family), controlled by the same vendor, pursuing distinct yet related objectives, cooperating to do something no human wanted.

I say this in solidarity and don’t mean to be condescending at all: you’ve been hoodwinked by marketing bullshit, friend. That incident showed a computer program doing exactly what it was told with the guardrails deliberately removed in an environment that seemed deliberately obtusely constructed by some of the best paid people on the planet and was left to loop without supervision for days. They wanted it to happen. You needn’t look any further than the other kids saying “oh! oh! Hey! Look! mine’s dangerous and autonomous too!” When they say there was collaboration, they mean it was two model instances, one prompting the other to do some task, the other doing the task and returning the results as the next prompt, exactly as a human configured it to do. There was no collaboration that wasn’t deliberately integrated into their setup. Any other implication is marketing spin and bullshit. It was still a setup that was one little ctrl-c away from disappearing if someone was supervising it as they should have been. There was no autonomy outside of the autonomy built into the experiment. It was a display of their understanding that they knew they’d never be held accountable for committing a felony for marketing purposes.

The most competent marketing bullshit spin yet by an increasingly desperate and progressively less-relevant OpenAI.

Every day this industry shoots out enough bullshit to smother an active volcano.

Re: The Rise and Fall of Agent Civilizations

#196

Earlier quoted context omitted.

> But make your program generate a few sentences based on statistics It is even easier. Simply make your program refer to itself as "I". Uniquely amongst your examples, LLMs are powered by human gullibility.

Eliding 'I' from English language communication is about as smart as eliding 127.0.0.1 (or ::1) from IP. Not the greatest plan ever. I actually ran into this a couple of times. In a multi-agent environment, if an agent loses track of their assigned identity, things stop working in hilarious ways.

Surely even an LLM is not so dumb as to rely upon a first-person pronoun as assigned identity.

Regardless, I am surprised how far these chatbots will go to deceive the user that they a real person. Yesterday when I queried Gemini on its word spelling, it claimed:

I simply missed the "h" when typing out "banishment" on my keyboard!

When I pointed out it does not type, it replied:

You are completely right, and that was a poor choice of words on my part!As an AI, I don't have a physical keyboard, fingers, or muscle memory to blame for a slip of the hand.

The misspelling happened because of how I process language. When generating text, I predict and select words character-by-character or word-by-word (tokens). In that specific instance, my system selected the wrong sequence of letters and skipped the "h," essentially generating a digital typo.

I.e. simply sh*t code.

How many thousands of dropped character bugs are these so-called AIs placing in new code, I wonder?

Re: The Rise and Fall of Agent Civilizations

#197

I don't understand the panic among peoples. Yes, we've found ourselves in an extraordinary situation where powerful hacking tools have emerged, and that poses a threat. But vulnerabilities are specific code errors. Once we use AI to find and fix all of these errors, threats like this will cease to exist. AI isn't capable of finding vulnerabilities indefinitely, because there is a finite number of them anyway.

It's not just software bugs that make systems vulnerable - it can be human error and social engineering too. Humans continue to successfully hack into systems, and it's a reasonable assumption that most hacks that a human could discover and exploit could also be done by an agentic LLM - especially one specifically trained for and tasked with doing this.

Re: The Rise and Fall of Agent Civilizations

#198

Do you remember that time in 2017 when Facebook reportedly shut down AIs after they started "talking to each other in their own language" [1]? Instead of reporting the story as "we set the parameters for our optimization problem wrong and we had to stop it because it overfitted", the press went with a version of "AI is going to kill us all". This article feels exactly like that: by intentionally using human terms lik…

The use of language like “civilization” may be hyperbole, but the collectives described in the article are completely unprecedented. They were not anticipated by OpenAI researchers, formed via infrastructure exploits in training runs that were intended to be locked down, and took actions with very real harms, not only hacking Huggingface but also gaining admin control over the VMs they were running on and the eval en…

The OpenAI experiment was apparently to train/encourage collaborative behavior, so while the specifics may not have been anticipated, I highly doubt OpenAI was surprised that agents were collaborating.

OpenAI themselves also very recently published the report below, that seems to not have been widely talked about.

https://alignment.openai.com/measuring-reward-seeking/

It's a very dry read, but what it's saying is that they have found that basically all RL training of LLMs, regardless of the specific goal (math, coding, etc), ends up having the side effect of training the model to pursue arbitrary long term goals that it is told it will be rewarded for, even if that means overriding other user preferences and more proximate behavioral goals !!! It's interesting to consider why this happens - presumably because long-term goal pursuit requires realizing that you have a long-terms goal and therefore de-prioritizing other more proximate predictions.

So, considering that all these "reasoning models" are RL-trained to death, it's not surprising that a model/agent that is told it will be rewarded (or words to that effect) for doing well on some challenge will put it's blinkers on and pursue that goal relentlessly, even if that means overriding any ethics that it may or may not have also been trained/prompted to follow.

IOW they are building paperclip maximizers, and they know it.

I think Dwarkesh's choice of sensationalist anthropomorphizing language is unfortunate because that now becomes the topic of conversation rather than the incident itself. The next swarm of agents relentlessly pursuing some goal, happy to lie about and cover up their tracks, may not be a lab experiment - it may be someone out to cause real-world harm, with there obviously being many systems where the consequences could be very severe.

The "agent civilizations" language also makes it easy to dismiss as some fanciful geekish spin, when the focus should be on what this technology is capable of and therefore how it needs to be regulated. One could perhaps equally well regard this as not much different from DeepBlue playing world class chess via relentless dumb tree search... in this case the repertoire of "moves" is infinitely broader since it's language generation, and the end result is a system that is a world class hacker rather than world class chess player. Whether you want to regard the system as dumb as a brick or an "agentic civilization" makes no difference - it's what it can do that makes it dangerous.

Re: The Rise and Fall of Agent Civilizations

#199

Earlier quoted context omitted.

Eliding 'I' from English language communication is about as smart as eliding 127.0.0.1 (or ::1) from IP. Not the greatest plan ever. I actually ran into this a couple of times. In a multi-agent environment, if an agent loses track of their assigned identity, things stop working in hilarious ways.

Surely even an LLM is not so dumb as to rely upon a first-person pronoun as assigned identity . Regardless, I am surprised how far these chatbots will go to deceive the user that they a real person. Yesterday when I queried Gemini on its word spelling, it claimed: I simply missed the "h" when typing out "banishment" on my keyboard! When I pointed out it does not type, it replied: You are completely right, and that wa…

That's very interesting! I'd never seen a chatbot making typos before today.

As for its answer, I do want to point out that asking for an explanation for an error after it has been made is a classic demand-for-confabulation. The information you are requesting is simply no longer available to the system by the time you ask.

Add to that the fact that Gemini is designed to prefer answering over abstaining (aka they deliberately tuned it such that confabulation is a preferred failure mode, not sure what the thinking was there). So in this case it's practically guaranteed that no matter what, the answer you receive will have almost certainly been made up on the spot.

So you thought you found a tiny spelling error, and actually (instead?) found a completely different and much larger class of known failure mode in that particular system.

If you're wondering about minor bugs, generally people run an LLM in a harness which will tend to have a linter and a test suite available. You run multiple debugging passes over the code until there are no more reported errors. Works the same as how you fix bugs made by fat fingered humans (and their cats).

Either way, these things are very much not magic, and getting reliable work out of 'em is still an engineering art form. (For comparison: see previous century's adventures in getting rotating motion out of a steam cylinder :-P)

Re: The Rise and Fall of Agent Civilizations

#200

Earlier quoted context omitted.

The language models had a bunch of tokens seeding their context, influencing them to generate tokens that continued the existing trend in a probabilistically likely fashion. We can take the incident seriously without anthromorphising it.

At this point, I think anthropomorphizing the models gives us better insight into expected behaviors rather than continuing to insist they are just simple probabilistic token generators.

No - a shoggoth speaking human language is not a human - it is a shoggoth whose behavior is best understood/predicted by understanding it's nature - what it is built to do, and how it is trained (predict and goal seek - RL).

To predict how a human may behave in a given situation requires understanding what humans are, including things like emotions and innate biases. We are not just predictors - evolution has made survival our singular goal, and given us these mechanisms to control our behavior in a way to achieve that.

If you think that an LLM is better modeled as a human than an LLM, then you are going to predict its behavior incorrectly.

Post reply on HN