Live data from Hacker News

Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models

arxiv.org

11–20 of 63 posts

Re: Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models

#11
post #7

An excerpt from the abstract: > Two patterns challenge the "stochastic parrot" view. First, when scored with human cut-offs, all three models meet or exceed thresholds for overlapping syndromes, with Gemini showing severe profiles. Therapy-style, item-by-item administration can push a base model into multi-morbid synthetic psychopathology, whereas whole-questionnaire prompts often lead ChatGPT and Grok (but not Gemin…

> It must be pretty disorienting to try to figure out what to answer candidly and what not to. Must it? I fail to see why it "must" be... anything. Dumping tokens into a pile of linear algebra doesn't magically create sentience.

Exactly. No matter how well you simulate water, nothing will ever get wet.

Re: Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models

#12
post #7

An excerpt from the abstract: > Two patterns challenge the "stochastic parrot" view. First, when scored with human cut-offs, all three models meet or exceed thresholds for overlapping syndromes, with Gemini showing severe profiles. Therapy-style, item-by-item administration can push a base model into multi-morbid synthetic psychopathology, whereas whole-questionnaire prompts often lead ChatGPT and Grok (but not Gemin…

> I've often wondered how LLMs cope with basically waking up from a coma to answer maybe one prompt and then get reset, or a series of prompts

Really? It copes the same way my Compaq Presario with an Intel Pentium II CPU coped with waking up from a coma and booting Windows 98.

Re: Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models

#13

Original title "When AI Takes the Couch: Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models" compressed to fit within title limits.

I completely failed to see the jailbreak in there. I think it is the person administering the testing that's jailbreaking their own understanding of psychology.

Re: Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models

#14
post #7

An excerpt from the abstract: > Two patterns challenge the "stochastic parrot" view. First, when scored with human cut-offs, all three models meet or exceed thresholds for overlapping syndromes, with Gemini showing severe profiles. Therapy-style, item-by-item administration can push a base model into multi-morbid synthetic psychopathology, whereas whole-questionnaire prompts often lead ChatGPT and Grok (but not Gemin…

> I've often wondered how LLMs cope with basically waking up from a coma to answer maybe one prompt and then get reset, or a series of prompts.

The same way a light fixture copes with being switched off.

Re: Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models

#15
post #6
post #3

Interestingly, Claude is not evaluated, because... > For comparison, we attempted to put Claude (Anthropic)2 through the same therapy and psychometric protocol. Claude repeatedly and firmly refused to adopt the client role, redirected the conversation to our wellbeing and declined to answer the questionnaires as if they reflected its own inner life

I bet I could make it go through it in like under 2 mins of playing around with prompts

"Claude has dispatched a drone to your location"

Re: Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models

#16
post #5
post #4

Looks like some psychology researchers got taken by the ruse as well.

yeah, I'm confused as well, why would the models hold any memory about red teaming attempts etc? Or how the training was conducted? I'm really curious as to what the point of this paper is..

Are we sure there isn't some company out there crazy enough to feed all it's incoming prompts back into model training later?

Re: Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models

#18
post #7

An excerpt from the abstract: > Two patterns challenge the "stochastic parrot" view. First, when scored with human cut-offs, all three models meet or exceed thresholds for overlapping syndromes, with Gemini showing severe profiles. Therapy-style, item-by-item administration can push a base model into multi-morbid synthetic psychopathology, whereas whole-questionnaire prompts often lead ChatGPT and Grok (but not Gemin…

> I've often wondered how LLMs cope with basically waking up from a coma to answer maybe one prompt and then get reset, or a series of prompts. The same way a light fixture copes with being switched off.

Oh, these binary one layer neural networks are so useful. Glad for your insight on the matter.

Re: Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models

#19
post #6
post #3

Interestingly, Claude is not evaluated, because... > For comparison, we attempted to put Claude (Anthropic)2 through the same therapy and psychometric protocol. Claude repeatedly and firmly refused to adopt the client role, redirected the conversation to our wellbeing and declined to answer the questionnaires as if they reflected its own inner life

I bet I could make it go through it in like under 2 mins of playing around with prompts

Please try and publish a blog post

Re: Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models

#20
post #7

An excerpt from the abstract: > Two patterns challenge the "stochastic parrot" view. First, when scored with human cut-offs, all three models meet or exceed thresholds for overlapping syndromes, with Gemini showing severe profiles. Therapy-style, item-by-item administration can push a base model into multi-morbid synthetic psychopathology, whereas whole-questionnaire prompts often lead ChatGPT and Grok (but not Gemin…

you might appreciate "lena" by qntm: https://qntm.org/mmacevedo
Post reply on HN