Live data from Hacker News

Large language models often know when they are being evaluated

arxiv.org

41–50 of 138 posts

Re: Large language models often know when they are being evaluated

#41
post #4

Just like they "know" English. "know" is quite an anthropomorphization. As long as an LLM will be able to describe what an evaluation is (why wouldn't it?) there's a reasonable expectation to distinguish/recognize/match patterns for evaluations. But to say they "know" is plenty of (unnecessary) steps ahead.

But do you know what it means to know?

I'm only being slightly sarcastic. Sentience is a scale. A worm has less than a mouse, a mouse has less than a dog, and a dog less than a human.

Sure, we can reset LLMs at will, but give them memory and continuity, and they definitely do not score zero on the sentience scale.

Re: Large language models often know when they are being evaluated

#42
post #4

Just like they "know" English. "know" is quite an anthropomorphization. As long as an LLM will be able to describe what an evaluation is (why wouldn't it?) there's a reasonable expectation to distinguish/recognize/match patterns for evaluations. But to say they "know" is plenty of (unnecessary) steps ahead.

But do you know what it means to know? I'm only being slightly sarcastic. Sentience is a scale. A worm has less than a mouse, a mouse has less than a dog, and a dog less than a human. Sure, we can reset LLMs at will, but give them memory and continuity, and they definitely do not score zero on the sentience scale.

If I set an LLM in a room by itself what does it do?

Re: Large language models often know when they are being evaluated

#43
post #40

Earlier quoted context omitted.

I agree with your point except for scientific papers. Let's push ourselves to use precise, non-shorthand or hand waving in technical papers and publications, yes? If not there, of all places, then where?

"Know" doesn't have any rigorous precisely-defined senses to be used! Asking for it not to be used colloquially is the same as asking for it never to be used at all. I mean - people have been saying stuff like "grep knows whether it's writing to stdout" for decades. In the context of talking about computer programs, that usage for "know" is the established/only usage, so it's hard to imagine any typical HN reader see…

colloquial use of "know" implies anthropomorphisation. Arguing that usign "knowing" in the title and "awarness" and "superhuman" in the abstract is just colloquial for "matching" is splitting hairs to an absurd degree.

Re: Large language models often know when they are being evaluated

#44

"...advanced reasoning models like Gemini 2.5 Pro and Claude-3.7-Sonnet (Thinking) can occasionally identify the specific benchmark origin of transcripts (including SWEBench, GAIA, and MMLU), indicating evaluation-awareness via memorization of known benchmarks from training data. Although such occurrences are rare, we note that because our evaluation datasets are derived from public benchmarks, memorization could pla…

[deleted]

Re: Large language models often know when they are being evaluated

#45
post #4

Just like they "know" English. "know" is quite an anthropomorphization. As long as an LLM will be able to describe what an evaluation is (why wouldn't it?) there's a reasonable expectation to distinguish/recognize/match patterns for evaluations. But to say they "know" is plenty of (unnecessary) steps ahead.

Communication is to vibration as knowledge is to resonance (?). From the sound of one hand clapping to the secret name of Ra.

Re: Large language models often know when they are being evaluated

#46

Earlier quoted context omitted.

But do you know what it means to know? I'm only being slightly sarcastic. Sentience is a scale. A worm has less than a mouse, a mouse has less than a dog, and a dog less than a human. Sure, we can reset LLMs at will, but give them memory and continuity, and they definitely do not score zero on the sentience scale.

If I set an LLM in a room by itself what does it do?

Yes, that's my fall back as well. If it receives zero instructions, will it take any action?

Re: Large language models often know when they are being evaluated

#47
post #39

Earlier quoted context omitted.

How would you prefer to describe this result then?

A term like knowing is fine if it is used in the abstract and then redefined more precisely in the paper. It isn't. Worse they start adding terms like scheming, pretending, awareness, and on and on. At this point you might as well take the model home and introduce it to your parents as your new life partner.

>A term like knowing is fine if it is used in the abstract and then redefined more precisely in the paper.

Sounds like a purely academic exercise.

Is there any genuine uncertainty about what the term "knowing" means in this context, in practice?

Can you name 2 distinct plausible definitions of "knowing", such that it would matter for the subject at hand which of those 2 definitions they're using?

Re: Large language models often know when they are being evaluated

#48
post #4

Just like they "know" English. "know" is quite an anthropomorphization. As long as an LLM will be able to describe what an evaluation is (why wouldn't it?) there's a reasonable expectation to distinguish/recognize/match patterns for evaluations. But to say they "know" is plenty of (unnecessary) steps ahead.

Communication is to vibration as knowledge is to resonance (?). From the sound of one hand clapping to the secret name of Ra.

I resonate with this vibe

Re: Large language models often know when they are being evaluated

#49
post #21

Earlier quoted context omitted.

How would you prefer to describe this result then?

One could say, for instance… A pattern matching algorithm detects when patterns match.

That's not what's going on here? The algorithms aren't being given any pattern of "being evaluated" / "not being evaluated", as far as I can tell. They're doing it zero-shot.

Put it another way: Why is this distinction important? We use the word "knowing" with humans. But one could also argue that humans are pattern-matchers! Why, specifically, wouldn't "knowing" apply to LLMs? What are the minimal changes one could make to existing LLM systems such that you'd be happy if the word "knowing" was applied to them?

Post reply on HN