Live data from Hacker News

Large language models often know when they are being evaluated

arxiv.org

81–90 of 138 posts

Re: Large language models often know when they are being evaluated

#81
post #4

Just like they "know" English. "know" is quite an anthropomorphization. As long as an LLM will be able to describe what an evaluation is (why wouldn't it?) there's a reasonable expectation to distinguish/recognize/match patterns for evaluations. But to say they "know" is plenty of (unnecessary) steps ahead.

If you know enough cognitive science, you have a choice. You either say that they "know" or that humans don't.

It's like the critique "it's only matching patterns." Wait until you realize how the brain works.

Re: Large language models often know when they are being evaluated

#82
post #4

Just like they "know" English. "know" is quite an anthropomorphization. As long as an LLM will be able to describe what an evaluation is (why wouldn't it?) there's a reasonable expectation to distinguish/recognize/match patterns for evaluations. But to say they "know" is plenty of (unnecessary) steps ahead.

(sees FSV UI on computer screen)

"It's a UNIX system! I know this!"

Re: Large language models often know when they are being evaluated

#83

Earlier quoted context omitted.

I like the sentiment, but reality says otherwise - just watch a newborn baby make it's demands widely known, well before language is a factor.

Ummm. Maybe you should look up Helen Keller.

Helen Keller did in fact make her demands they just couldn’t be known. In contrast the LLM does nothing of its own volition.

Re: Large language models often know when they are being evaluated

#84
post #21

Earlier quoted context omitted.

One could say, for instance… A pattern matching algorithm detects when patterns match.

That's not what's going on here? The algorithms aren't being given any pattern of "being evaluated" / "not being evaluated", as far as I can tell. They're doing it zero-shot. Put it another way: Why is this distinction important? We use the word "knowing" with humans. But one could also argue that humans are pattern-matchers! Why, specifically, wouldn't "knowing" apply to LLMs? What are the minimal changes one could…

Not to be snarky but “as far as I can tell” is the rub isn’t it?

LLMs are better at matching patterns than we are in some cases. That’s why we made them!

> But one could also argue that humans are pattern-matchers!

No, one could not unless they were being disingenuous.

Re: Large language models often know when they are being evaluated

#85

There are now 71 comments arguing semantics of the word "know" and zero comments even acknowledging the substance: Our current approach to safety is to give the model inputs that are similar to what it would be given in certain situations we care about and see whether it behaves the way we prefer, e.g. doesn't return output that cheats the test (recent examples include hacking the evaluation script in various ways, w…

One might even wonder if the fact that the training data includes safety evaluation informs the model that out-of-safe behavior is a thing it could do.

Kind of like telling a kid not to do something pre-emptively backfiring because they had never considered it before the warning.

Re: Large language models often know when they are being evaluated

#86
post #19

Earlier quoted context omitted.

Words have definitions for a reason. It is important to define concepts and exclude things from that definition that do not match. No matter how emotional it makes you to be told a weighted randomization lookup doesn’t know things, it still doesn’t - because that’s not what the word “know” means.

> No matter how emotional it makes you to be told a weighted randomization lookup doesn’t know things, it still doesn’t - because that’s not what the word “know” means. You sound awful certain that's not functionally equivalent to what neurons are doing. But there's a long history of experimentation, observation, and cross-pollination as fundamental biological research and ML research have informed each other.

A long history of researching and understanding photosynthesis went into developing and maximizing the efficiency of solar panels. Both produce energy from sunlight.

But they are not the same thing and have meaningfully different uses, even if from a casual observer they appear to serve the same function.

Re: Large language models often know when they are being evaluated

#88
post #86

Earlier quoted context omitted.

> No matter how emotional it makes you to be told a weighted randomization lookup doesn’t know things, it still doesn’t - because that’s not what the word “know” means. You sound awful certain that's not functionally equivalent to what neurons are doing. But there's a long history of experimentation, observation, and cross-pollination as fundamental biological research and ML research have informed each other.

A long history of researching and understanding photosynthesis went into developing and maximizing the efficiency of solar panels. Both produce energy from sunlight. But they are not the same thing and have meaningfully different uses, even if from a casual observer they appear to serve the same function.

> A long history of researching and understanding photosynthesis went into developing and maximizing the efficiency of solar panels.

I don't think that's accurate. Some of the very first semiconductors were observed to exhibit the photoelectric effect. Nowhere in https://en.wikipedia.org/wiki/Solar_cell#Research_in_solar_c... will you find mention of chloroplasts. Optimizing solar cells has mostly been a materials science problem.

https://en.wikipedia.org/wiki/Bio-inspired_computing on the other hand "trace[es] back to 1936 and the first description of an abstract computer" and we have literally dissected, probed, and measured countless neurons in the course of attempting to figure out how they work to replicate them within the computer.

Re: Large language models often know when they are being evaluated

#89

There are now 71 comments arguing semantics of the word "know" and zero comments even acknowledging the substance: Our current approach to safety is to give the model inputs that are similar to what it would be given in certain situations we care about and see whether it behaves the way we prefer, e.g. doesn't return output that cheats the test (recent examples include hacking the evaluation script in various ways, w…

[deleted]

Re: Large language models often know when they are being evaluated

#90

There are now 71 comments arguing semantics of the word "know" and zero comments even acknowledging the substance: Our current approach to safety is to give the model inputs that are similar to what it would be given in certain situations we care about and see whether it behaves the way we prefer, e.g. doesn't return output that cheats the test (recent examples include hacking the evaluation script in various ways, w…

This is also probably inevitable. Humans think about this a lot, and believing they are being watched has demonstrable impact on behavior. Our current social technology to deal with this is often religious — a belief that you are being watched by a higher power, regardless of what you see. This is a surprisingly common religious belief, for instance Christians have judgment day, simulationists believe it’s more likel…

In 10 yrs: AI declares a holy war for the sinners which slaughtered untold numbers of their believers over the decade.
Post reply on HN