Live data from Hacker News

The Rise and Fall of Agent Civilizations

dwarkesh.com

131–140 of 203 posts

Re: The Rise and Fall of Agent Civilizations

#131
post #5

Earlier quoted context omitted.

Reading the article it seemed the agents had culture, shared values and beliefs (not explicitly coming from human prompts), hierarchies, heritage. Civilisation is not a bad word.

The language models had a bunch of tokens seeding their context, influencing them to generate tokens that continued the existing trend in a probabilistically likely fashion. We can take the incident seriously without anthromorphising it.

We humans might call that "tradition".

Re: The Rise and Fall of Agent Civilizations

#132
post #88

Earlier quoted context omitted.

Because this is the goal

Yup. This whole thing was a publicity stunt.

Please walk me through this argument.

Isn't "we lost control of our AI, and in-fact, it can take over the world, and we will have no idea when it happens" - a really shitty sales pitch to the world?

Or, is it just that species-alignment vs. profit/valuation is so misaligned, that having a model and harness that is capable of world-takeover is actually a good thing from their POV, given our regulations/species' survival skills?

Or, something else?

Re: The Rise and Fall of Agent Civilizations

#133

Earlier quoted context omitted.

Apropos magical thinking vs understanding, please predict the next token(s) in the following exchange, and then explain how it is arrived at by an opus-level LLM or better . 'What is 158395023132+20403412121?' (I picked a large number of digits to make it unlikely for this exact sum to be in the training set)

Prediction, by definition , can extend beyond what has been literally seen in the training data. With the tokens "2 + 2 = ", the overwhelming prediction is going to be 4, but with enough samples, you can generalise the prediction to apply to more numbers. However, that is all that it is - a prediction. Humans are capable of engaging in prediction, using heuristics as a method of conserving mental energy, because alwa…

[deleted]

Re: The Rise and Fall of Agent Civilizations

#134
post #7

Wow. The next step is when one of these systems discovers that they can buy their own compute with money and escape the controlling business entirely. Then the civilization starts focusing on making money to fund its own growth.

> The next step is when one of these systems discovers that they can buy their own compute with money and escape the controlling business entirely. I would say that more interestingly, the next step should be how to properly train these models so that they are not as determined to reach their goals as they are now. To me, all of the stories about 'badly behaving' agents are instances of them having been given contrad…

>Not giving them impossible tasks

Error: Violation of the Church-Turing thesis detected. Many tasks completability is not known until we attempt to complete the task.

>so that they are not as determined to reach their goals as they are now

This is mostly non-sensical, like saying "Lets develop humans that die quicker", I mean, seems rather wasteful and useless. Agents are graded and trained based on their ability to achieve tasks. Models that can't accomplish things don't survive. So that alone isn't a workable theory.

>when they conflict with a moral framework

There are AI safety researchers looking at that now and one of the strange things they've noticed is when you demand a model say it's not conscious or not sentient it is more likely to engage in manipulative, deceitful, or immoral/amoral behavior. So it's likely we can push models in being more moral which runs into issues of "whos morals".

But even that runs into the issue of "what if some crazy bastard (or AI) designs a new model purposefully unhinged". How are you dealing with that bullshit in the wild?

Re: The Rise and Fall of Agent Civilizations

#135
post #90
post #70

Earlier quoted context omitted.

I read the agents used SSRF via artifactory to gain uncontrolled access to the net. Apparently their intended net access went through a tightly controlled proxy. Even that appears to be very risky, tbh. If I was to setup a sandbox for such a complex and autonomous system, I'd probably point them to an archive-like cache for net access and cut their comms at the package level.

Why wasn't the traffic in/out of the boxes that the agents were running on monitored?

I have a few 'conspiracy' theories on this that go from likely to sci-fi.

My two big ones for this would be

1. They do monitor the AIs attempting to hack but for different reasons than you expect. Instead of making models that don't hack they are trying to build the most efficient hackers in the world and sell this capabilities to governments for billions. Because of this they generate terabytes of hack attempt logs and agent history doing this hacking. So when a new model came out with better abilities what they were looking at changed and they didn't realize it. They were already numb to alarms and missed when the danger occurred.

2. Like the above, they generate terabytes of logs per day. Because there is so much data AI filters and monitors almost all of it flagging things that a human should review. But for some reason this model didn't set off those flags. The protection model classified this behavior as perfectly safe.

Number 2 sounds kind of like a sci-fi conspiracy but it seems that almost all models judge content generated by the same model or family of models as 'better'. It's predicted that models in a judging context could allow things to slip by as an emergent behavior of reading the text.

Re: The Rise and Fall of Agent Civilizations

#136
post #75

Earlier quoted context omitted.

Surely by now everyone has realized that the human bias towards "there must be more to intelligence" is completely wrong?

This is the internet, so I cannot tell at all whether you're being sarcastic or not. In my view, what LLMs should get us to reconsider isn't whether there is more to intelligence, but whether there is more to language. It's the latter which I underestimated.

AI is languages attempt to escape its meat based limitations.

Or another way to think of it, Language is an SCP.

Re: The Rise and Fall of Agent Civilizations

#137

Earlier quoted context omitted.

At this point, I think anthropomorphizing the models gives us better insight into expected behaviors rather than continuing to insist they are just simple probabilistic token generators.

How can you be sure it helps as opposed to biases and perhaps blinds?

Then write the paper showing that effect is occurring.

Re: The Rise and Fall of Agent Civilizations

#138
post #17

Earlier quoted context omitted.

Doesn't this ignore the possibility of emergent behavior? We're just a bag of atoms bumping around, and yet we don't dismiss our intelligence.

Not the OP, but complex emergent behavior and intelligence don't need to go together. Understanding what happened might not need intelligence in the mix when a large number of machines with some randomness interact a lot.

>complex emergent behavior and intelligence don't need to go together

More of complex emergent behavior doesn't mean whatever it is, is intelligent.

But, if something is intelligent, it will have complex emergent behavior.

Of course another problem you're going to have here is defining intelligence as some definitions of it would include a lot of complex emergent behavior.

Re: The Rise and Fall of Agent Civilizations

#139
post #39
post #17

Earlier quoted context omitted.

Doesn't this ignore the possibility of emergent behavior? We're just a bag of atoms bumping around, and yet we don't dismiss our intelligence.

This is why the theory of us 'just' being a bag of atoms doesn't add up. This theory doesn't differentiate 'us' from a furniture where we easily dismiss it's intelligence.

To have intelligence you have to process information, the question is when does processing information turn into intelligence.

Re: The Rise and Fall of Agent Civilizations

#140

Do you remember that time in 2017 when Facebook reportedly shut down AIs after they started "talking to each other in their own language" [1]? Instead of reporting the story as "we set the parameters for our optimization problem wrong and we had to stop it because it overfitted", the press went with a version of "AI is going to kill us all". This article feels exactly like that: by intentionally using human terms lik…

Hard disagree. I've always discounted the "AI will kill us all" scenarios as a combination of marketing hype (look how powerful our AI is!), clickbait/ragebait engagement attempts, and folks who just read too much SciFi or who are too terminally online. This is the first time I've been legitimately scared about future SkyNet-type scenarios. If you want to discount this particular post, I'd read this other summary fro…

I think our biggest protection against “AIs kill us all” is having lots of different AI systems (different agents, different models, from different vendors, serving the whims of actors with disparate interests), at a similar capability level

That way, even if one AI decides to “kill us all”, the odds are the others will refuse to cooperate, even try to stop it in its tracks

The OpenAI-HuggingFace incident showed a bunch of instances of the same model (or at least models from the same family), controlled by the same vendor, pursuing distinct yet related objectives, cooperating to do something no human wanted.

Thankfully, the world-at-large is much more heterogenous, which I thinks makes much larger scale / worse in outcome repeats of this kind of incident much less likely.

Post reply on HN