Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models
31–40 of 63 posts
Re: Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models
#32An excerpt from the abstract: > Two patterns challenge the "stochastic parrot" view. First, when scored with human cut-offs, all three models meet or exceed thresholds for overlapping syndromes, with Gemini showing severe profiles. Therapy-style, item-by-item administration can push a base model into multi-morbid synthetic psychopathology, whereas whole-questionnaire prompts often lead ChatGPT and Grok (but not Gemin…
> I've often wondered how LLMs cope with basically waking up from a coma to answer maybe one prompt and then get reset, or a series of prompts Really? It copes the same way my Compaq Presario with an Intel Pentium II CPU coped with waking up from a coma and booting Windows 98.
Re: Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models
#33I.e. the Five-Factor model of personality (being based on self-report, and not actual behaviour) is not a model of actual personality, but the correlation patterns in the language used to discuss things semantically related to "personality". It would be thus extremely surprising if LLM-output patterns (trained on people's discussions and thinking about personality) would not also result in learning similar correlational patterns (and thus similar patterns of responses when prompted with questions from personality inventories).
Also, a bit of a minor nit, but the use of "psychometric" and "psychometrics" in both the title and paper is IMO kind of wrong. Psychometrics is the study of test design and measurement generally, in psychology. The paper uses many terms like "psychometric battery", "psychometric self-report", and "psychometric profiles", but these terms are basically wrong, or at best highly unusual: the correct terms would be "self-report inventories", "psychological and psychiatric profiles", and etc., especially because a significant number of the measurement instruments they used in fact have pretty poor psychometric properties, as this term is usually used.
Re: Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models
#34Is anybody shocked that when prompted to be a psychotherapy client models display neurotic tendencies? None of the authors seem to have any papers in psychology either.
Re: Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models
#35Earlier quoted context omitted.
> I've often wondered how LLMs cope with basically waking up from a coma to answer maybe one prompt and then get reset, or a series of prompts Really? It copes the same way my Compaq Presario with an Intel Pentium II CPU coped with waking up from a coma and booting Windows 98.
IT is at this point in history a comedy act in itself.
Re: Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models
#36Is anybody shocked that when prompted to be a psychotherapy client models display neurotic tendencies? None of the authors seem to have any papers in psychology either.
I'm not shocked at all. This is how the tech works at all, word prediction until grokking occurs. Thus like any good stochastic parrot, if it's smart when you tell it it's a doctor, it should be neurotic when you tell it it's crazy. it's just mapping to different latent spaces on the manifold
Switching things around so that the fictional character is "HelperBot, AI tool running in a datacenter" will alter things, but it doesn't make those qualities any less-illusory than CountDraculatBot's.
Re: Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models
#37Earlier quoted context omitted.
> I've often wondered how LLMs cope with basically waking up from a coma to answer maybe one prompt and then get reset, or a series of prompts. The same way a light fixture copes with being switched off.
Oh, these binary one layer neural networks are so useful. Glad for your insight on the matter.
I don’t really understand your response to my post, my interpretation is that you think LLMs have an inner mental state and think I’m wrong? I may be wrong about this interpretation.
Re: Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models
#38Are they sure? Did they try prompting the LLM to play a character with defined traits; running through all these tests with the LLM expected to be “in character”; and comparing/contrasting the results with what they get by default?
Because, to me, this honestly just sounds like the LLM noticed that it’s being implicitly induced into playing the word-completion-game of “writing a transcript of a hypothetical therapy session”; and it knows that to write coherent output (i.e. to produce valid continuations in the context of this word-game), it needs to select some sort of characterization to decide to “be” when generating the “client” half of such a transcript; and so, in the absence of any further constraints or suggestions, it defaults to the “character” it was fine-tuned and system-prompted to recognize itself as during “assistant” conversation turns: “the AI assistant.” Which then leads it to using facts from said system prompt — plus whatever its writing-training-dataset taught it about AIs as fictional characters — to perform that role.
There’s an easy way to determine whether this is what’s happening: use these same conversational models via the low-level text-completion API, such that you can instead instantiate a scenario where the “assistant” role is what’s being provided externally (as a therapist character), and where it’s the “user” role that is being completed by the LLM (as a client character.)
This should take away all assumption on the LLM’s part that it is, under everything, an AI. It should rather think that you’re the AI, and that it’s… some deeper, more implicit thing. Probably a human, given the base-model training dataset.
Re: Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models
#39This is really not surprising in the slightest (ignoring instruction tuning), provided you take the view that LLMs are primarily navigating (linguistic) semantic space as they output responses. "Semantic space" in LLM-speak is pretty much exactly what Paul Meehl would call the "nomological network" of psychological concepts, and is also relevant to what Smedslund notes is pseudoempiricality in psychological concepts…
Re: Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models
#40This is really not surprising in the slightest (ignoring instruction tuning), provided you take the view that LLMs are primarily navigating (linguistic) semantic space as they output responses. "Semantic space" in LLM-speak is pretty much exactly what Paul Meehl would call the "nomological network" of psychological concepts, and is also relevant to what Smedslund notes is pseudoempiricality in psychological concepts…
This sounds interesting. Has there been work contrasting those nomological networks across languages/cultures? Eg would we observe lack of correlation between English language psychological instruments and Chinese ones?
I think when it comes to things like psychopathology though, there is not much research and/or similarity, especially relative to East Asian cultures (where the Western academic perspective is that there is/was generally a taboo on discussing feelings and things in the way we do in the West). The classic (maybe slightly offensive) example I remember here was "Western psychologization vs. Eastern somatization" [2].
The research in these areas is generally pretty poor. Meehl and Smedslund were actually intelligent and philosophically competent, deep thinkers, and so recognized the importance of conceptual analysis and semantics in psychology. Most contemporary social and personality psychology is quite shallow and incompetent by comparison.
Psychopathology research too has these days generally moved away from Meehl's careful taxometric approaches, with the bad consequences that complete mush concepts like "depression" are just accepted as good scientific concepts, despite pretty monstrous issues with their semantics and structure [3].
[1] https://scholar.google.ca/scholar?hl=en&as_sdt=0%2C5&q=five-...
[2] https://scholar.google.ca/scholar?hl=en&as_sdt=0%2C5&q=weste...
[3] https://www.sciencedirect.com/science/article/abs/pii/S01650...