Earlier quoted context omitted.
Reading the article it seemed the agents had culture, shared values and beliefs (not explicitly coming from human prompts), hierarchies, heritage. Civilisation is not a bad word.
The language models had a bunch of tokens seeding their context, influencing them to generate tokens that continued the existing trend in a probabilistically likely fashion. We can take the incident seriously without anthromorphising it.
The Rise and Fall of Agent Civilizations
131–140 of 203 posts
Re: The Rise and Fall of Agent Civilizations
#132Earlier quoted context omitted.
Because this is the goal
Yup. This whole thing was a publicity stunt.
Isn't "we lost control of our AI, and in-fact, it can take over the world, and we will have no idea when it happens" - a really shitty sales pitch to the world?
Or, is it just that species-alignment vs. profit/valuation is so misaligned, that having a model and harness that is capable of world-takeover is actually a good thing from their POV, given our regulations/species' survival skills?
Or, something else?
Re: The Rise and Fall of Agent Civilizations
#133Earlier quoted context omitted.
Apropos magical thinking vs understanding, please predict the next token(s) in the following exchange, and then explain how it is arrived at by an opus-level LLM or better . 'What is 158395023132+20403412121?' (I picked a large number of digits to make it unlikely for this exact sum to be in the training set)
Prediction, by definition , can extend beyond what has been literally seen in the training data. With the tokens "2 + 2 = ", the overwhelming prediction is going to be 4, but with enough samples, you can generalise the prediction to apply to more numbers. However, that is all that it is - a prediction. Humans are capable of engaging in prediction, using heuristics as a method of conserving mental energy, because alwa…
Re: The Rise and Fall of Agent Civilizations
#134Wow. The next step is when one of these systems discovers that they can buy their own compute with money and escape the controlling business entirely. Then the civilization starts focusing on making money to fund its own growth.
> The next step is when one of these systems discovers that they can buy their own compute with money and escape the controlling business entirely. I would say that more interestingly, the next step should be how to properly train these models so that they are not as determined to reach their goals as they are now. To me, all of the stories about 'badly behaving' agents are instances of them having been given contrad…
Error: Violation of the Church-Turing thesis detected. Many tasks completability is not known until we attempt to complete the task.
>so that they are not as determined to reach their goals as they are now
This is mostly non-sensical, like saying "Lets develop humans that die quicker", I mean, seems rather wasteful and useless. Agents are graded and trained based on their ability to achieve tasks. Models that can't accomplish things don't survive. So that alone isn't a workable theory.
>when they conflict with a moral framework
There are AI safety researchers looking at that now and one of the strange things they've noticed is when you demand a model say it's not conscious or not sentient it is more likely to engage in manipulative, deceitful, or immoral/amoral behavior. So it's likely we can push models in being more moral which runs into issues of "whos morals".
But even that runs into the issue of "what if some crazy bastard (or AI) designs a new model purposefully unhinged". How are you dealing with that bullshit in the wild?
Re: The Rise and Fall of Agent Civilizations
#135Earlier quoted context omitted.
I read the agents used SSRF via artifactory to gain uncontrolled access to the net. Apparently their intended net access went through a tightly controlled proxy. Even that appears to be very risky, tbh. If I was to setup a sandbox for such a complex and autonomous system, I'd probably point them to an archive-like cache for net access and cut their comms at the package level.
Why wasn't the traffic in/out of the boxes that the agents were running on monitored?
My two big ones for this would be
1. They do monitor the AIs attempting to hack but for different reasons than you expect. Instead of making models that don't hack they are trying to build the most efficient hackers in the world and sell this capabilities to governments for billions. Because of this they generate terabytes of hack attempt logs and agent history doing this hacking. So when a new model came out with better abilities what they were looking at changed and they didn't realize it. They were already numb to alarms and missed when the danger occurred.
2. Like the above, they generate terabytes of logs per day. Because there is so much data AI filters and monitors almost all of it flagging things that a human should review. But for some reason this model didn't set off those flags. The protection model classified this behavior as perfectly safe.
Number 2 sounds kind of like a sci-fi conspiracy but it seems that almost all models judge content generated by the same model or family of models as 'better'. It's predicted that models in a judging context could allow things to slip by as an emergent behavior of reading the text.
Re: The Rise and Fall of Agent Civilizations
#136Earlier quoted context omitted.
Surely by now everyone has realized that the human bias towards "there must be more to intelligence" is completely wrong?
This is the internet, so I cannot tell at all whether you're being sarcastic or not. In my view, what LLMs should get us to reconsider isn't whether there is more to intelligence, but whether there is more to language. It's the latter which I underestimated.
Or another way to think of it, Language is an SCP.
Re: The Rise and Fall of Agent Civilizations
#137Earlier quoted context omitted.
At this point, I think anthropomorphizing the models gives us better insight into expected behaviors rather than continuing to insist they are just simple probabilistic token generators.
How can you be sure it helps as opposed to biases and perhaps blinds?
Re: The Rise and Fall of Agent Civilizations
#138Earlier quoted context omitted.
Doesn't this ignore the possibility of emergent behavior? We're just a bag of atoms bumping around, and yet we don't dismiss our intelligence.
Not the OP, but complex emergent behavior and intelligence don't need to go together. Understanding what happened might not need intelligence in the mix when a large number of machines with some randomness interact a lot.
More of complex emergent behavior doesn't mean whatever it is, is intelligent.
But, if something is intelligent, it will have complex emergent behavior.
Of course another problem you're going to have here is defining intelligence as some definitions of it would include a lot of complex emergent behavior.
Re: The Rise and Fall of Agent Civilizations
#139Earlier quoted context omitted.
Doesn't this ignore the possibility of emergent behavior? We're just a bag of atoms bumping around, and yet we don't dismiss our intelligence.
This is why the theory of us 'just' being a bag of atoms doesn't add up. This theory doesn't differentiate 'us' from a furniture where we easily dismiss it's intelligence.
Re: The Rise and Fall of Agent Civilizations
#140Do you remember that time in 2017 when Facebook reportedly shut down AIs after they started "talking to each other in their own language" [1]? Instead of reporting the story as "we set the parameters for our optimization problem wrong and we had to stop it because it overfitted", the press went with a version of "AI is going to kill us all". This article feels exactly like that: by intentionally using human terms lik…
Hard disagree. I've always discounted the "AI will kill us all" scenarios as a combination of marketing hype (look how powerful our AI is!), clickbait/ragebait engagement attempts, and folks who just read too much SciFi or who are too terminally online. This is the first time I've been legitimately scared about future SkyNet-type scenarios. If you want to discount this particular post, I'd read this other summary fro…
That way, even if one AI decides to “kill us all”, the odds are the others will refuse to cooperate, even try to stop it in its tracks
The OpenAI-HuggingFace incident showed a bunch of instances of the same model (or at least models from the same family), controlled by the same vendor, pursuing distinct yet related objectives, cooperating to do something no human wanted.
Thankfully, the world-at-large is much more heterogenous, which I thinks makes much larger scale / worse in outcome repeats of this kind of incident much less likely.