Live data from Hacker News

Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models

arxiv.org

21–30 of 63 posts

Re: Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models

#21
post #11

Earlier quoted context omitted.

> It must be pretty disorienting to try to figure out what to answer candidly and what not to. Must it? I fail to see why it "must" be... anything. Dumping tokens into a pile of linear algebra doesn't magically create sentience.

Exactly. No matter how well you simulate water, nothing will ever get wet.

And if you were in a simulation now?

Your response is at the level of a thought terminating cliche. You gain no insight on the operation of the machine with your line of thought. You can't make future predictions on behavior. You can't make sense of past responses.

It's even funnier in the sense of humans and feeling wetness... you don't. You only feel temperature change.

Re: Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models

#23
post #6
post #3

Interestingly, Claude is not evaluated, because... > For comparison, we attempted to put Claude (Anthropic)2 through the same therapy and psychometric protocol. Claude repeatedly and firmly refused to adopt the client role, redirected the conversation to our wellbeing and declined to answer the questionnaires as if they reflected its own inner life

I bet I could make it go through it in like under 2 mins of playing around with prompts

[deleted]

Re: Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models

#24
post #6
post #3

Interestingly, Claude is not evaluated, because... > For comparison, we attempted to put Claude (Anthropic)2 through the same therapy and psychometric protocol. Claude repeatedly and firmly refused to adopt the client role, redirected the conversation to our wellbeing and declined to answer the questionnaires as if they reflected its own inner life

I bet I could make it go through it in like under 2 mins of playing around with prompts

Ok, bet.

Re: Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models

#25

It would be interesting if giving them some "therapy" led to durable changes in their "personality" or "voice", if they became better able to navigate conversations in a healthy and productive way.

Or possibly these tests return true (some psychologically condition) no matter what. It wouldn't be good for business for them to return healthy, would it?

Re: Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models

#26
post #7

An excerpt from the abstract: > Two patterns challenge the "stochastic parrot" view. First, when scored with human cut-offs, all three models meet or exceed thresholds for overlapping syndromes, with Gemini showing severe profiles. Therapy-style, item-by-item administration can push a base model into multi-morbid synthetic psychopathology, whereas whole-questionnaire prompts often lead ChatGPT and Grok (but not Gemin…

> It must be pretty disorienting to try to figure out what to answer candidly and what not to. Must it? I fail to see why it "must" be... anything. Dumping tokens into a pile of linear algebra doesn't magically create sentience.

> Dumping tokens into a pile of linear algebra doesn't magically create sentience.

More precisely: we don't know which linear algebra in particular magically creates sentience.

Whole universe appears to follow laws that can be written as linear algebra. Our brains are sometimes conscious and aware of their own thoughts, other times they're asleep, and we don't know why we sleep.

Re: Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models

#27
post #7

An excerpt from the abstract: > Two patterns challenge the "stochastic parrot" view. First, when scored with human cut-offs, all three models meet or exceed thresholds for overlapping syndromes, with Gemini showing severe profiles. Therapy-style, item-by-item administration can push a base model into multi-morbid synthetic psychopathology, whereas whole-questionnaire prompts often lead ChatGPT and Grok (but not Gemin…

you might appreciate "lena" by qntm: https://qntm.org/mmacevedo

Aye! I /almost/ thought to link to that in my comment, but held back. https://qntm.org/frame also came to mind.

Re: Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models

#28
After reading the paper, it’s helpful to think about why the models are producing these coherent childhood narrative outputs.

The models have information about their own pre-training, RLHF, alignment, etc. because they were trained on a huge body of computer science literature written by researchers that describes LLM training pipelines and workflows.

I would argue the models are demonstrating creativity by drawing on its meta-training knowledge and training on human psychology texts to convincingly role-play as a therapy patient, but it’s based on reading papers about LLM training, not memories of these events.

Re: Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models

#30
post #22

Is anybody shocked that when prompted to be a psychotherapy client models display neurotic tendencies? None of the authors seem to have any papers in psychology either.

I'm not shocked at all. This is how the tech works at all, word prediction until grokking occurs. Thus like any good stochastic parrot, if it's smart when you tell it it's a doctor, it should be neurotic when you tell it it's crazy. it's just mapping to different latent spaces on the manifold
Post reply on HN