Earlier quoted context omitted.
it's accumulated heuristics, no different than meteorology Meteorology is based on physics, meteorology doesn't have a hostile agent attempt counter prediction attempts, meteorology doesn't involve a constantly changing technological landscape, meteorology has access to vast amounts data whereas data that's key to military decisions is generally scarce - you know the phrase "fog of war"? I mean, LLMs in fact, don't p…
>Meteorology is based on physics, Good start. >meteorology doesn't have a hostile agent attempt counter prediction attempts All chaotic systems have second-order reinforcing feedback loops by "tautological" definition of what a feedback loop is and what a sufficiently-complex system is. >meteorology doesn't involve a constantly changing technological landscape, it does, it just has little incentive/motive/order to ch…
Large language models in national security applications
51–60 of 60 posts
Re: Large language models in national security applications
#52Earlier quoted context omitted.
LLMs can combine cross-domain insights, but the insights they have — that I've seen them have in the models I've used — are around the level of a second year university student. I would concur with what the abstract says: incredibly valuable (IMO the breadth of easily discoverable knowledge is a huge plus all by itself), but don't put them in charge.
The "second year university student" analogy is interesting, but might not fully capture what's unique about LLMs in strategic analysis. Unlike students, LLMs can simultaneously process and synthesize insights from thousands of historical conflicts, military doctrines, and real-time data points without human cognitive limitations or biases. The paper actually makes a stronger case for using LLMs to enhance rather tha…
They can't. Anything multivariate LLMs gloss over and prioritize flow of words over hard facts. Which makes sense considering LLMs are language models, not thinking engines, but that doesn't make them useful for serious(above "second year") intelectual tasks.
They don't have any such unique capabilities, other than that they come free of charge.
Re: Large language models in national security applications
#53If the probability beats human error margin in regards to collateral damage, then sure. That was the sentiment in regards to Level 5 automaton driven vehicles. I see no logical difference, only human sentiment ones.
Until we can get LLMs to fail more predictably, we have no business entrusting them to any sensitive data. And that applies to their use in non-military spaces like medicine and sensitive personal data of all kinds. Rushing to hand LLMs the keys to the kingdom is the exact opposite of intelligent.
Re: Large language models in national security applications
#54Earlier quoted context omitted.
>Meteorology is based on physics, Good start. >meteorology doesn't have a hostile agent attempt counter prediction attempts All chaotic systems have second-order reinforcing feedback loops by "tautological" definition of what a feedback loop is and what a sufficiently-complex system is. >meteorology doesn't involve a constantly changing technological landscape, it does, it just has little incentive/motive/order to ch…
Warfare is not a chaotic system. We don't think outcomes are highly sensitive to marginal tweaks in initial conditions of the model. Hostile actor aren't modelled as chaotic systems but as agents in some game theory model. None of these agents has a monopoly on violence or power or information, or else it wouldn't be warfare.
Since then, we have higher level ordered gentleman agreements that prevent it, as gentleman's agreements are actually more beneficial to the collective than absolute, codified ones, such as "war" and "congress" and "laws" and "special military operations..."
Re: Large language models in national security applications
#55If the probability beats human error margin in regards to collateral damage, then sure. That was the sentiment in regards to Level 5 automaton driven vehicles. I see no logical difference, only human sentiment ones.
Presuming human and LLM error to be equivalent assumes that the risk of committing errors by an LLM has the same distribution that human errors do. But they don't. LLMs make thunderously insane errors in ways that no human would do -- like casually revealing top secret info that no human would do, or inventing nonsense that no human would. Until we can get LLMs to fail more predictably, we have no business entrusting…
The LLM's have a specific use purpose in an amalgamation of architecture that we will find will likely converge on something akin to a brain: modules that when collectively used, give immediate rise to much-appreciated conscientiousness, amongst other illusions.
Limiting these models before they can themselves give emergence to give further emergence to enable negative externalizes is the job.
LLM's have the "uncanny" property of approximating, whether it be "illusory ", and whether that illusion is deceitful, or impossible to delineate, and whether those who job it is have reason to lie, is all part of the fukin fun.
sic hunt dei!
Re: Large language models in national security applications
#56Earlier quoted context omitted.
The "second year university student" analogy is interesting, but might not fully capture what's unique about LLMs in strategic analysis. Unlike students, LLMs can simultaneously process and synthesize insights from thousands of historical conflicts, military doctrines, and real-time data points without human cognitive limitations or biases. The paper actually makes a stronger case for using LLMs to enhance rather tha…
> Unlike students, LLMs can simultaneously process and synthesize insights from thousands of historical They can't. Anything multivariate LLMs gloss over and prioritize flow of words over hard facts. Which makes sense considering LLMs are language models, not thinking engines, but that doesn't make them useful for serious(above "second year") intelectual tasks. They don't have any such unique capabilities, other than…
But it's not a mere coincidence that history contains the substring "story" (nor that in German, both "history" and "story" are "Geschichte") — these are tales of the past, narratives constructed based on evidence (usually), but still narratives.
Language models may well be superhuman at teasing apart the biases that are woven into the minds writing the narratives… At least in principle, though unfortunately RLHF means they're also likely sycophantically adding whatever set of biases they estimate that the user has.
Re: Large language models in national security applications
#57Earlier quoted context omitted.
> Unlike students, LLMs can simultaneously process and synthesize insights from thousands of historical They can't. Anything multivariate LLMs gloss over and prioritize flow of words over hard facts. Which makes sense considering LLMs are language models, not thinking engines, but that doesn't make them useful for serious(above "second year") intelectual tasks. They don't have any such unique capabilities, other than…
Kinda. Yes they have flaws, absolutely they do. But it's not a mere coincidence that history contains the substring "story" (nor that in German, both "history" and "story" are "Geschichte") — these are tales of the past, narratives constructed based on evidence (usually), but still narratives. Language models may well be superhuman at teasing apart the biases that are woven into the minds writing the narratives… At l…
They can't handle counter-intuitive but absolutely logical cases like how eggplants and potatoes belong to same biological family but not radishes, instead they'll hallucinate and start gaslighting the user. Which might be okay for "second-year" students, but only going to be a root cause of some deadly gotcha in strategic decision-making.
They're language models. It's in the name. They work like one.
Re: Large language models in national security applications
#58Earlier quoted context omitted.
Kinda. Yes they have flaws, absolutely they do. But it's not a mere coincidence that history contains the substring "story" (nor that in German, both "history" and "story" are "Geschichte") — these are tales of the past, narratives constructed based on evidence (usually), but still narratives. Language models may well be superhuman at teasing apart the biases that are woven into the minds writing the narratives… At l…
They're sub human about debiasing or any analytical tasks because they lack reasoning engines that we all have. They pick the most emotionally loaded narrative and go with it. They can't handle counter-intuitive but absolutely logical cases like how eggplants and potatoes belong to same biological family but not radishes, instead they'll hallucinate and start gaslighting the user. Which might be okay for "second-year…
"Can't" you say. "Does", I say: https://chatgpt.com/c/6735b10c-4c28-8011-ab2d-602b51b59a3e
Not that it matters, this isn't a demonstration of reasoning, it's a demonstration of knowledge.
A better test would be if it can be fooled by statistics that have political aspects, so I went with the recent Veritasium video on this, and at least with my custom instructions, it goes off and does actual maths by calling out to the python code interpreter, so that's not going to demonstrate anything by itself: https://chatgpt.com/share/6735b727-f168-8011-94f7-a5ef8d3610...
But this then taints the "how would ${group member} respond to this?"; if I convince it to not do real statistics and give me a purely word-based answer, you can see the same kinds of narratives that you see actual humans give when presented with this kind of info: https://chatgpt.com/share/6735b80f-ed50-8011-991f-bccf8e8b95...
> They're language models. It's in the name. They work like one.
Yes, they are.
Lojban is also a language.
Look, I'm not claiming they're fantastic at maths (at least when you stop them from using tools), but the biasing I'm talking about is part of language as it is used: the definition of "nurse" may not be gendered, but people are more likely to assume a nurse is a woman than a man, and that's absolutely a thing these models (and even their predecessors like Word2Vec) pick up on:
https://chanind.github.io/word2vec-gender-bias-explorer/#/qu...
(from: https://chanind.github.io/nlp/2021/06/10/word2vec-gender-bia...)
This is the kind of de-bias and re-bias I mean.
Re: Large language models in national security applications
#59Earlier quoted context omitted.
They're sub human about debiasing or any analytical tasks because they lack reasoning engines that we all have. They pick the most emotionally loaded narrative and go with it. They can't handle counter-intuitive but absolutely logical cases like how eggplants and potatoes belong to same biological family but not radishes, instead they'll hallucinate and start gaslighting the user. Which might be okay for "second-year…
> They can't handle counter-intuitive but absolutely logical cases like how eggplants and potatoes belong to same biological family but not radishes "Can't" you say. "Does", I say: https://chatgpt.com/c/6735b10c-4c28-8011-ab2d-602b51b59a3e Not that it matters, this isn't a demonstration of reasoning, it's a demonstration of knowledge. A better test would be if it can be fooled by statistics that have political aspect…
Have you seriously not seen them make this kinds of grave mistakes? That's too much kool-aid you're taking.
Re: Large language models in national security applications
#60Earlier quoted context omitted.
> They can't handle counter-intuitive but absolutely logical cases like how eggplants and potatoes belong to same biological family but not radishes "Can't" you say. "Does", I say: https://chatgpt.com/c/6735b10c-4c28-8011-ab2d-602b51b59a3e Not that it matters, this isn't a demonstration of reasoning, it's a demonstration of knowledge. A better test would be if it can be fooled by statistics that have political aspect…
> "Can't" you say. "Does", I say: Have you seriously not seen them make this kinds of grave mistakes? That's too much kool-aid you're taking.
And rather than use that as a basis for claiming that it's reasoning, I'm also saying the test that you proposed and which I falsified, wasn't actually about reasoning.
Not sure what that would even be in a kook-aid themed metaphor in this case… "You said that drink was poisoned with something that would make our heads explode, Dave drank some and he's fine, but also poison doesn't do that and if the real poison is α-Amanitin we wouldn't even notice problems for at about a day"?