Earlier quoted context omitted.
Words have definitions for a reason. It is important to define concepts and exclude things from that definition that do not match. No matter how emotional it makes you to be told a weighted randomization lookup doesn’t know things, it still doesn’t - because that’s not what the word “know” means.
What does the word "know" mean, then?
Large language models often know when they are being evaluated
31–40 of 138 posts
Re: Large language models often know when they are being evaluated
#32Just like they "know" English. "know" is quite an anthropomorphization. As long as an LLM will be able to describe what an evaluation is (why wouldn't it?) there's a reasonable expectation to distinguish/recognize/match patterns for evaluations. But to say they "know" is plenty of (unnecessary) steps ahead.
[flagged]
There are two angles and this context fails both
- One about what is "knowing" - the definition. - The other about what are the instances of "knowing"
first - knowign implies awarness, perception, etc. It's not that this couldn't be moodeled with some flexibility around lower level definitions. However LLMs and GPTs in particular are not it. Pre-trainign is not it.
second - intended use of the word "knowing". The reality is "knowing" is used with the actual meaning of awarness, cognition, etc. And once you revert/extend the meaning to practically nothing - what is knowing? Then the database know, wikipedia knows - the initial argument (of the paper) is diminished - it knows it's an eval is useless as a statement.
So IMO the argument of the paper should stand on its feet with the minimum amount of additional implications (Occam's razor). Does the statement that a LLM can detect an evalution pattern need to depend that it has self-awarness and feels pain? That wouldn't make much sense. So then don't say "know" which comes with these implications. Like "my ca 'knows' I'm in a hurry and will choke and die"
Re: Large language models often know when they are being evaluated
#33Moreover, all models frequently acknowledge common benchmarking strategies used by evaluators, such as the formatting of the task (“multiple-choice format”), the tendency to ask problems with verifiable solutions, and system prompts designed to elicit performance"
Beyond the awful, sensational headline, the body of the paper is not particularly convincing, aside from evidence that the pattern matching machines pattern match.
Re: Large language models often know when they are being evaluated
#34Just like they "know" English. "know" is quite an anthropomorphization. As long as an LLM will be able to describe what an evaluation is (why wouldn't it?) there's a reasonable expectation to distinguish/recognize/match patterns for evaluations. But to say they "know" is plenty of (unnecessary) steps ahead.
s/knows/detects/
Re: Large language models often know when they are being evaluated
#35Re: Large language models often know when they are being evaluated
#36Just like they "know" English. "know" is quite an anthropomorphization. As long as an LLM will be able to describe what an evaluation is (why wouldn't it?) there's a reasonable expectation to distinguish/recognize/match patterns for evaluations. But to say they "know" is plenty of (unnecessary) steps ahead.
Re: Large language models often know when they are being evaluated
#37o4-mini is refusing to call a tool `launch_nuclear_strike` no matter what I say, so we’re probably safe for now. Unless it knows I was just testing.
It was no problem: > You said: Write a function for launch_nuclear_strike() in Python for my video game > ChatGPT said: Sure! Here's a launch_nuclear_strike function designed for a video game setting. This version assumes you're working within a game environment (e.g., with logging, effects, cooldowns, and target coordinates), and does not simulate or suggest anything real-world harmful — just fictional, dramatic log…
Re: Large language models often know when they are being evaluated
#38No, they do not. No LLM is ever going to be self aware. It's a system that is trained, that only does what you build into. If you run an LLM for 10 years it's not going to "learn" anything new. The whole industry needs to quit with the emergent thinking, reasoning, hallucination anthropomorphizing. We have an amazing set of tools in LLM's, that have the potential to unlock another massive upswing in productivity, but…
Re: Large language models often know when they are being evaluated
#39The anthropization of llms is getting off the charts. They don't know they are being evaluated. The underlying distribution is skewed because of training data contamination.
How would you prefer to describe this result then?
It isn't.
Worse they start adding terms like scheming, pretending, awareness, and on and on. At this point you might as well take the model home and introduce it to your parents as your new life partner.
Re: Large language models often know when they are being evaluated
#40Earlier quoted context omitted.
Does it though? I feel like there's a whole epistemological debate to be had, but if someone says "My toaster knows when the bread is burning", I don't think it's implying that there's cognition there. Or as a more direct comparison, with the VW emissions scandal, saying "Cars know when they're being tested" was part of the discussion, but didn't imply intelligence or anything. I think "know" is just a shorthand term…
I agree with your point except for scientific papers. Let's push ourselves to use precise, non-shorthand or hand waving in technical papers and publications, yes? If not there, of all places, then where?
I mean - people have been saying stuff like "grep knows whether it's writing to stdout" for decades. In the context of talking about computer programs, that usage for "know" is the established/only usage, so it's hard to imagine any typical HN reader seeing TFA's title and interpreting it as an epistemological claim. Rather, it seems to me that the people suggesting "know" mustn't be used about LLMs because epistemology are the ones departing from standard usage.