Live data from Hacker News

Large language models often know when they are being evaluated

arxiv.org

51–60 of 138 posts

Re: Large language models often know when they are being evaluated

#51
post #4

Just like they "know" English. "know" is quite an anthropomorphization. As long as an LLM will be able to describe what an evaluation is (why wouldn't it?) there's a reasonable expectation to distinguish/recognize/match patterns for evaluations. But to say they "know" is plenty of (unnecessary) steps ahead.

But do you know what it means to know? I'm only being slightly sarcastic. Sentience is a scale. A worm has less than a mouse, a mouse has less than a dog, and a dog less than a human. Sure, we can reset LLMs at will, but give them memory and continuity, and they definitely do not score zero on the sentience scale.

It probably scores about the same as a calculator, which I’d say is zero.

Re: Large language models often know when they are being evaluated

#52

Earlier quoted context omitted.

If I set an LLM in a room by itself what does it do?

Yes, that's my fall back as well. If it receives zero instructions, will it take any action?

Helen Keller famously said that before she had language (the first word of which was “water”) she had nothing, a void, and the minute she had language, “the whole world came rushing in.”

Perhaps we are not so very different?

Re: Large language models often know when they are being evaluated

#53
post #43
post #40

Earlier quoted context omitted.

"Know" doesn't have any rigorous precisely-defined senses to be used! Asking for it not to be used colloquially is the same as asking for it never to be used at all. I mean - people have been saying stuff like "grep knows whether it's writing to stdout" for decades. In the context of talking about computer programs, that usage for "know" is the established/only usage, so it's hard to imagine any typical HN reader see…

colloquial use of "know" implies anthropomorphisation. Arguing that usign "knowing" in the title and "awarness" and "superhuman" in the abstract is just colloquial for "matching" is splitting hairs to an absurd degree.

You missed the substance of my comment. Certainly the title is anthropomorphism - and anthropomorphism is a rhetorical device, not a scientific claim. The reader can understand that TFA means it non-rigorously, because there is no rigorous thing for it to mean.

As such, to me the complaint behind this thread falls into the category of "I know exactly what TFA meant but I want to argue about how it was phrased", which is definitely not my favorite part of the HN comment taxonomy.

Re: Large language models often know when they are being evaluated

#54
post #32

Earlier quoted context omitted.

[flagged]

"Knowing" needs not exist outside of human invention. In fact that's the point - it only matters in relation to humans. You can choose whatever definition you want, but the reality is that, once you chose a non-standard definition the argument becomes meaningless outside of the scope of your definition. There are two angles and this context fails both - One about what is "knowing" - the definition. - The other about…

>"Knowing" needs not exist outside of human invention. In fact that's the point

It doesn't need to, I never said it needed to. That is my point. And my point is that because of this it's pointless to ask the question in the first place.

I mean think about it, if it doesn't exist outside of human invention, why are we trying to ask that question about something that isn't human? An LLM?

Re: Large language models often know when they are being evaluated

#56

Earlier quoted context omitted.

But do you know what it means to know? I'm only being slightly sarcastic. Sentience is a scale. A worm has less than a mouse, a mouse has less than a dog, and a dog less than a human. Sure, we can reset LLMs at will, but give them memory and continuity, and they definitely do not score zero on the sentience scale.

If I set an LLM in a room by itself what does it do?

Is the LLM allowed to do anything without prompting? Or is it effectively disabled? This is more a question of the setup than of sentience.

Re: Large language models often know when they are being evaluated

#57
post #53
post #43

Earlier quoted context omitted.

colloquial use of "know" implies anthropomorphisation. Arguing that usign "knowing" in the title and "awarness" and "superhuman" in the abstract is just colloquial for "matching" is splitting hairs to an absurd degree.

You missed the substance of my comment. Certainly the title is anthropomorphism - and anthropomorphism is a rhetorical device, not a scientific claim. The reader can understand that TFA means it non-rigorously, because there is no rigorous thing for it to mean. As such, to me the complaint behind this thread falls into the category of "I know exactly what TFA meant but I want to argue about how it was phrased", which…

I see. Thanks for clarifying. I did want to argue about how it was phrased and what is alluding to. Implying increased risk from "knowing" the eval regime is roughly as weak as the definition of "knowing". It can be equaly a measure of general detection capability, as it can about evaluation incapability - i.e. unlikely news worthy, unless it reached top HN because of the "know" in the title.

Re: Large language models often know when they are being evaluated

#58
post #57
post #53

Earlier quoted context omitted.

You missed the substance of my comment. Certainly the title is anthropomorphism - and anthropomorphism is a rhetorical device, not a scientific claim. The reader can understand that TFA means it non-rigorously, because there is no rigorous thing for it to mean. As such, to me the complaint behind this thread falls into the category of "I know exactly what TFA meant but I want to argue about how it was phrased", which…

I see. Thanks for clarifying. I did want to argue about how it was phrased and what is alluding to. Implying increased risk from "knowing" the eval regime is roughly as weak as the definition of "knowing". It can be equaly a measure of general detection capability, as it can about evaluation incapability - i.e. unlikely news worthy, unless it reached top HN because of the "know" in the title.

Thanks for replying - I kind of follow you but I only skimmed the paper. To be clear I was more responding to the replies about cognition, than to what you said about the eval regime.

Incidentally I think you might be misreading the paper's use of "superhuman"? I assume it's being used to mean "at a higher rate than the human control group", not (ironically) in the colloquial "amazing!" sense.

Re: Large language models often know when they are being evaluated

#59
post #39

Earlier quoted context omitted.

A term like knowing is fine if it is used in the abstract and then redefined more precisely in the paper. It isn't. Worse they start adding terms like scheming, pretending, awareness, and on and on. At this point you might as well take the model home and introduce it to your parents as your new life partner.

>A term like knowing is fine if it is used in the abstract and then redefined more precisely in the paper. Sounds like a purely academic exercise. Is there any genuine uncertainty about what the term "knowing" means in this context, in practice? Can you name 2 distinct plausible definitions of "knowing", such that it would matter for the subject at hand which of those 2 definitions they're using?

> Sounds like a purely academic exercise.

Well, yes. It’s an academic research paper (I assume since it’s submitted to arXiv) and to be submitted to academic journals/conferences/etc., so it’s a fairly reasonable critique of the authors/the paper.

Re: Large language models often know when they are being evaluated

#60
post #3

The anthropization of llms is getting off the charts. They don't know they are being evaluated. The underlying distribution is skewed because of training data contamination.

> The anthropization of llms is getting off the charts.

What's wrong with that? If it quacks like a duck... it's just a complex pile of organic chemistry, ducks aren't real because the concept of "a duck" is wrong.

I honestly believe there is a degree of sentience in LLMs. Sure, they're not sentient in the human sense, but if you define sentience as whatever humans have, then of course no other entity can be sentient.

Post reply on HN