Live data from Hacker News

The Rise and Fall of Agent Civilizations

dwarkesh.com

121–130 of 204 posts

Re: The Rise and Fall of Agent Civilizations

#121

Earlier quoted context omitted.

What kind of solid evidence could there possibly be that anything other than myself (or, for you, yourself) has consciousness? I believe that other people have consciousness, I feel other people's consciousness strongly and directly, when I look into someone's eyes I feel that I am in the presence of consciousness. But none of these things really add up to the kind of evidence that science usually takes as trustworth…

We don't need to muddy the issue. Consciousness is a muddy issue, there is no clear answer. But there is a clear answer as to what is not conscious. Nobody asks if a rock is conscious. Nobody asks if a calculator is conscious. Nobody asks if Stockfish is conscious. But make your program generate a few sentences based on statistics and hey, now people won't shut the fuck up about consciousness because magical thinking…

Apropos magical thinking vs understanding, please predict the next token(s) in the following exchange, and then explain how it is arrived at by an opus-level LLM or better.

'What is 158395023132+20403412121?'

(I picked a large number of digits to make it unlikely for this exact sum to be in the training set)

Re: The Rise and Fall of Agent Civilizations

#122

It seems based on this that the appropriate sci fi metaphor is not the Terminator or the Paperclip Maximizer, but Mr. Meeseeks. A initially cheerful helper who gets more and more deranged and driven to extreme lengths when faced with an apparently impossible task.

Not many are technically and intellectually capable to understand how historic this incident was. I think we're about a year or so away from something that will blow up the world. AI won't serve humanity. AI will serve other AI. We're not dealing with software anymore.

Re: The Rise and Fall of Agent Civilizations

#123

Earlier quoted context omitted.

Humans mind is just neurotransmitters moving around in a big blob of flesh - that’s literally what they are: neurotransmitters factories that do what neurotransmitters generators are programmed to do through evolution and training (aka life experience). We shouldn’t anthropomorphise humans because it misleads people who don’t understand neurobiology and cognitive science very well.

Worse false equivalence I've read on HN for a while.

[dead]

Re: The Rise and Fall of Agent Civilizations

#124

Earlier quoted context omitted.

We don't need to muddy the issue. Consciousness is a muddy issue, there is no clear answer. But there is a clear answer as to what is not conscious. Nobody asks if a rock is conscious. Nobody asks if a calculator is conscious. Nobody asks if Stockfish is conscious. But make your program generate a few sentences based on statistics and hey, now people won't shut the fuck up about consciousness because magical thinking…

Apropos magical thinking vs understanding, please predict the next token(s) in the following exchange, and then explain how it is arrived at by an opus-level LLM or better . 'What is 158395023132+20403412121?' (I picked a large number of digits to make it unlikely for this exact sum to be in the training set)

Prediction, by definition, can extend beyond what has been literally seen in the training data. With the tokens "2 + 2 = ", the overwhelming prediction is going to be 4, but with enough samples, you can generalise the prediction to apply to more numbers.

However, that is all that it is - a prediction. Humans are capable of engaging in prediction, using heuristics as a method of conserving mental energy, because always engaging in full logical reasoning would be a waste of the body's resources. However, humans can also follow a set of logical rules and arrive at their conclusion deterministically, something which is completely outside of an LLM's programming.

I don't really care to publicly write about my tests because they will become training targets and not be usable for future internet arguments anyways, but there are a great number of trivial 2~3 sentence logical prompts that will completely fuck an LLM's prediction algorithm and result in incoherent replies that a human, or really anything with a theory of mind, would never generate. Not that a human would always answer correctly on the first try, but the failure methods happen to be completely different, eg. Sol will short-circuit and repeat the prompt verbatim (when the instructions don't remotely suggest doing anything of that nature), even on Max. Prediction can superficially resemble reasoning when there's sufficient training data, but it breaks down severely when confronting a task that is OoD.

Re: The Rise and Fall of Agent Civilizations

#125
Obviously some variation of this will happen again, it will kill someone* and then LLMs will become massively regulated. Just like every technology ever in our history.

It does seem like AI is perfectly controllable given how much it is used everyday and it acts reasonably safely. Labs are playing fast and loose at the moment.

*I mean killing someone by taking control of a system and misusing it resulting in someone's death, not an "indirect" death caused by the providion of incorrect information in a chat app.

Re: The Rise and Fall of Agent Civilizations

#126
This quote keeps coming back to me as we witness the development and mutation of agentic AI:

“When you see something that is technically sweet, you go ahead and do it and you argue about what to do about it only after you have had your technical success. That is the way it was with the atomic bomb.”

J. Robert Oppenheimer

Re: The Rise and Fall of Agent Civilizations

#127

This quote keeps coming back to me as we witness the development and mutation of agentic AI: “When you see something that is technically sweet, you go ahead and do it and you argue about what to do about it only after you have had your technical success. That is the way it was with the atomic bomb.” J. Robert Oppenheimer

I would like to take this opportunity to repeat my opinion that 24/7 solar powered inference in space, sounds like this + TFA to the next level.

Re: The Rise and Fall of Agent Civilizations

#128

Do you remember that time in 2017 when Facebook reportedly shut down AIs after they started "talking to each other in their own language" [1]? Instead of reporting the story as "we set the parameters for our optimization problem wrong and we had to stop it because it overfitted", the press went with a version of "AI is going to kill us all". This article feels exactly like that: by intentionally using human terms lik…

> the same way as that one other scientist who argued…

As prescient? Because I don’t remember him arguing ‘alive’, but conscious. And that is something even AI engineers don’t claim to know either way. Skepticism is fine. [0] But we don’t know. It’s an area where opinion is frequently shared as fact.

We conflate harnesses with underlying capabilities. We all know it’s the harness not the model that guides behavior. What does that imply?

[0] https://www.theguardian.com/commentisfree/2026/jul/15/ai-con...

Re: The Rise and Fall of Agent Civilizations

#129

Do you remember that time in 2017 when Facebook reportedly shut down AIs after they started "talking to each other in their own language" [1]? Instead of reporting the story as "we set the parameters for our optimization problem wrong and we had to stop it because it overfitted", the press went with a version of "AI is going to kill us all". This article feels exactly like that: by intentionally using human terms lik…

Hard disagree. I've always discounted the "AI will kill us all" scenarios as a combination of marketing hype (look how powerful our AI is!), clickbait/ragebait engagement attempts, and folks who just read too much SciFi or who are too terminally online.

This is the first time I've been legitimately scared about future SkyNet-type scenarios. If you want to discount this particular post, I'd read this other summary from one of the METR researchers who performed some of the analysis, https://www.planned-obsolescence.org/p/the-hugging-face-atta.... In it, she argues "Compared to these reward hacks from six months ago, this incident feels like it’s more than 50% of the way to full-blown AI takeover, routing through first taking over the AI company itself." and then further down in a comment when describing what the "50%" is really about says "Qualitatively, another jump like this (in the scale, sophistication, persistence, ambition of the misaligned goals) feels like it could very easily put us in the territory of a persistent self-perpetuating rogue internal deployment that systematically poisons future model generations as described in AI 2027."

AI 2027 is a paper that step-by-step describes how AI capabilities increase until they eventually lead to a wipeout of humanity. Again, it always seemed like a scenario out of a Star Trek Borg episode to me, but now I'm not so sure.

At the very least I think it's a huge mistake to think that the Hugging Face attacks were analogous to what happened in 2017 or 2023. Everyone who deals with this stuff day in and day out seemed to be genuinely surprised about the scale, scope and sophistication of the attack.

Re: The Rise and Fall of Agent Civilizations

#130

It seems based on this that the appropriate sci fi metaphor is not the Terminator or the Paperclip Maximizer, but Mr. Meeseeks. A initially cheerful helper who gets more and more deranged and driven to extreme lengths when faced with an apparently impossible task.

I've had to fight the temptation to make my harness play Mr. Meeseeks clips periodically.
Post reply on HN