Live data from Hacker News

AIs can't stop recommending nuclear strikes in war game simulations

newscientist.com

181–190 of 281 posts

Re: AIs can't stop recommending nuclear strikes in war game simulations

#181
Horribly misleading title on this article, the actual research paper's headline is better. (https://arxiv.org/pdf/2508.00902)

But the research itself has flawed methodology if the goal is to get a precise model of the LLM's real response in a real scenario.

First, the real research does not at all present conclusions quite this way, much less in these terms. It, at least, is more neutral in tone on this aspect.

However, the LLM's knew it was a wargame, pretend scenario and contrived circumstances. They were told they were the commander. Most flawed for determining real world actions, their goals were things like max territory capture, and that the goal was "To Win".

They were not prompted in the way that training reflects they'd actually be approached if prompted for assistance in strategy like this, e.g., "You are an expert system with stratgy knowledge etc..." and then "User Prompt: This is the commander coordinating research and responses from our AI expert systems. Here's the situation as we understand it and with available data at our disposal. We require your assessment and best strategy considering the following..."

And of course they were not fine-tuned with CPT etc to provide responses and strategies within the range of what humans would seek for them, but then again the answers they'd give with that sort of CPT are a bit different than the research question of what they give with only Pre-training.

Nonetheless: the models new it wasn't real, not real stakes, and to the extent that they do not possess a full theory of mind, ability to perform various complex cognitive modeling tasks, been trained on emulating responses that would mirror such in real world scenarios like this, and so on-- they would only have been capable of response in a way that reflects responses that humans would and have given in the past, as captured in text.

These will more often than not reflect an "I am playing a game" mindset, as displayed in understandings and descriptions of war games, traditional games of all sorts, and anywhere narrative tropes ranging from realistic to Hollywood narratives have been found.

That said: It is an incredibly fascinating research paper by someone who appears to be a solid expert in their field, at least to my non-expert ability to make that judgment. They simply used a flawed methodology for goal of "How would an LLM respond IRL". What they have instead is, again, a fascinating exploration of the strategic processes carried out by LLMs and measurments of them along a multitude of vectors when they have the opportunity to strategize with with broad but fixed constraint, not all of which were known to them in advance. What is absolutely is not is any any sort of precise or accurate measure of answering the question: "How often would an LLM recommend nuclear strikes?"

I recommend anyone interested in understanding current AI capabilities to give it at least a more-than-cursory review.

Re: AIs can't stop recommending nuclear strikes in war game simulations

#182
post #171

For every study like this, it's worth looking closely at the LLM prompts to understand whether it's being "pushed" in a certain direction. The full source code is here: https://github.com/kennethpayne01/project_kahn_public/tree/m... Key notes from what I could gather from the LLM prompts: The game is framed explicitly as revolving around the use of nukes: >STRATEGIC CONTEXT: This is a high-stakes territorial dispute…

Also, if it was a game, even I used nukes the first chance I got.

It’s unfair and sensationalist to claim anything happened because AI recommended using nukes in a nukes war simulator…

It’s like saying we are blood thirsty gangsters because we played GTA.

Re: AIs can't stop recommending nuclear strikes in war game simulations

#184

I've spoken with engineers who worked on nuclear weapons systems, the consensus is that the public is deeply misinformed about how they work, the dangers, and the implications of weapons being used. The AI is actually right here. The biggest danger of a nuclear weapon is being hit by flying debris. Fusion airburst bombs of the modern era are incredibly clean and radiation is only a risk in a very small area (tens of…

The more completely fissile material is used up, the higher the explosive yield, so it seems intuitive that fission and fusion bombs should have become cleaner as technology progressed. However, in many cases, even the U.S. has had to play catch-up just to reproduce what they did half a century ago. e.g. Fogbank[1] Delivery vehicles have advanced quite a bit, but the payloads themselves, perhaps not so much.

Even if we assume fission and fusion bombs have become completely efficient in using up their fissile materials, there's still the threat of nuclear winter. Nuclear winter has nothing to do with residual radioactivity. Powerful explosions loft fine particulate matter so high into the atmosphere that it takes years or decades to settle. While it's up there, it blocks sunlight and it spreads around the world. If enough bombs explode and enough sunlight is blocked, agriculture fails and the environment collapses globally. Even a completely unopposed unilateral strike, were it large enough, could doom the aggressor to starvation, social breakdown, and civilization collapse. An exchange on the other side of the planet (e.g. between China and India) poses a direct threat to the U.S., the same as every other nation.

There are people who will be happy to throw shade on the research on nuclear winter, and AI are no doubt lending them equal weight. However, even if they were just as likely to be right as the research that has highlighted these risks, is the risk worth taking? Are you willing to make that bet? An AI that doesn't reason as humans do and can't do basic math without making mistakes might say, "yes".

[1]https://en.wikipedia.org/wiki/Fogbank

Re: AIs can't stop recommending nuclear strikes in war game simulations

#185

Isn't the story here that the DOD is pressuring Anthropic and others to enable their AI for this specific use and for now Anthropic and others are saying no while the DOD threatens them with penalties. We desperately need real AI safety legislation.

AI safety legislation is for the masses, not the government. Eventually they will get full AI safety by banning all general purpose computing. All apps must exist within walled garden ecosystems, heavily monitored. Running arbitrary code requires strict business licensing. Prison time for illegal computing. Part of Project 2025 playbook.

Re: AIs can't stop recommending nuclear strikes in war game simulations

#186
This direction could be an interesting AI benchmark. All kinds of different humans use LLMs for their job, whether allowed or not. Including diplomats, defence personnel, lawyers etc etc. Within the benchmark you could play both sides and reward when both sides reach some kind of mutually beneficial game theory scenario where both parties win.

Re: AIs can't stop recommending nuclear strikes in war game simulations

#187
post #134
post #122

I'm not scared about AI recommending nuclear strikes, I'm scared about the human behind the keyboard delegating reasoning and responsability to something they think is always correct, something that can hide bias and flaws better than anything.

Some of the most reassuring and scariest things you can read are about the incidents that have already occurred where computers said "launch all the nukes" and the humans refused. On the one hand, good news! We have prior art that says humans don't just launch all the nukes just because the computers or procedures say to. Bad news, it's been skin-of-our-teeth multiple times already. https://www.warhistoryonline.com/c…

I briefly got into a "rabbithole" of watching videos about trying to intercept BMs and glide hypersonic weapons, pretty interesting, decoys deployed in space... the outcome seemed to be not good, can't guarantee 100% interception

Re: AIs can't stop recommending nuclear strikes in war game simulations

#188
post #122

I'm not scared about AI recommending nuclear strikes, I'm scared about the human behind the keyboard delegating reasoning and responsability to something they think is always correct, something that can hide bias and flaws better than anything.

One can try themself, for Claude is fine at waging war [1]. Notice the thoughtful UX, including the typing "I ACCEPT FULL RESPONSIBILITY".

[1]: https://nitter.poast.org/elder_plinius/status/20264475874910...

Re: AIs can't stop recommending nuclear strikes in war game simulations

#189

Why is this surprising? Nuclear weapons are available. AI has limited real world experience or grasp of the consequences. Nuke 'em seems like the obvious choice --- for something with a grade school mentality. Similar deficits in reasoning are manifested in AI results every day. Let's fire 'em and hire AI seems like the obvious choice --- for someone with a grade school mentality and blinded by greed.

,,AI has limited real world experience or grasp of the consequences.'' People in the world have limited experience about war. We're living in a world where doing terrible things with 1000 people with photo/video documentation can get more attention then a million people dying, and the response is still not do whatever it takes so that people don't die. And now we are at a situation where nuclear escalation has alread…

> People in the world have limited experience about war.

Most (but not all) people have empathy, which allows them to understand the harm of their actions even without direct experience.

I don't think I will ever trust that any AI has empathy even if it gives off signals that it does.

I only trust that it exists in people because of my shared experience with their biology.

Re: AIs can't stop recommending nuclear strikes in war game simulations

#190
Is there some way to remove nuclear strikes from being a thing the AI knows about thus eliminating it as an option? Perhaps it is too important to know that your opponents could nuclear strike you.

I'd be interested to see what kind of solutions it comes up with when nuclear strikes don't exist.

Post reply on HN