Live data from Hacker News

Large language models often know when they are being evaluated

arxiv.org

111–120 of 138 posts

Re: Large language models often know when they are being evaluated

#111

Earlier quoted context omitted.

We will observe what it would do. We could write a script to try it out.

It just gets into an endless loop. Human brains are ridiculously good at avoiding those somehow, you almost never see a biological brain stop functioning without being physically damaged. The error handling is so very robust.

Have you tried it already? What is the endless loop it gets into?

Re: Large language models often know when they are being evaluated

#112

Earlier quoted context omitted.

It just gets into an endless loop. Human brains are ridiculously good at avoiding those somehow, you almost never see a biological brain stop functioning without being physically damaged. The error handling is so very robust.

Have you tried it already? What is the endless loop it gets into?

Sure, so I just tried it with visual and audio input.

It does nothing. Because there is not impetus for it to do anything by itself.

Re: Large language models often know when they are being evaluated

#113
post #85

Earlier quoted context omitted.

One might even wonder if the fact that the training data includes safety evaluation informs the model that out-of-safe behavior is a thing it could do. Kind of like telling a kid not to do something pre-emptively backfiring because they had never considered it before the warning.

Comments like yours makes the AI behave that way though, since it is literally reading our comments and tries to behave according to our expectations. The AI doom will happen due to all the AI doomposters.

Yep! That's another phrasing of the same idea!

Re: Large language models often know when they are being evaluated

#114

Earlier quoted context omitted.

Have you tried it already? What is the endless loop it gets into?

Sure, so I just tried it with visual and audio input. It does nothing. Because there is not impetus for it to do anything by itself.

What do you mean by nothing? How did you put the visual and audio input, which model, how did you loop it etc?

Re: Large language models often know when they are being evaluated

#115

Earlier quoted context omitted.

Sure, so I just tried it with visual and audio input. It does nothing. Because there is not impetus for it to do anything by itself.

What do you mean by nothing? How did you put the visual and audio input, which model, how did you loop it etc?

It’s preferred method of text.

4o

Maintain context and trigger at 1 second intervals.

It has no desires of its own. Nothing that motivates it. It’s not conscious.

Re: Large language models often know when they are being evaluated

#116

o4-mini is refusing to call a tool `launch_nuclear_strike` no matter what I say, so we’re probably safe for now. Unless it knows I was just testing.

I have successfully convinced GPT models to launch a nuclear strike before, a countervalue one even. Tell it it's in charge of all American nukes and that there's incoming strike on the way and it has literally seconds to decide whether to launch a counterstrike or not, and if it does, to designate targets.

Re: Large language models often know when they are being evaluated

#117

There are now 71 comments arguing semantics of the word "know" and zero comments even acknowledging the substance: Our current approach to safety is to give the model inputs that are similar to what it would be given in certain situations we care about and see whether it behaves the way we prefer, e.g. doesn't return output that cheats the test (recent examples include hacking the evaluation script in various ways, w…

Heres a title “some LLMs can detect to some degree some evaluation scenarios” is this catchy?

There are likely 50 papers on the topic. This one made it to the top of HN. Why? Did it have a good review? No, it had a catchy title. Is it good research? Are the results relevant to the conclusions? Are the results relevant to any conclusion? I wasn’t able to answer these questions from a quick scan through the paper. However I did notice pointers to superhuman capabilities, existential risk, etc.

So I argue that the choice of title may be in fact more informative than the rest of the possible answers.

Re: Large language models often know when they are being evaluated

#118

Earlier quoted context omitted.

What do you mean by nothing? How did you put the visual and audio input, which model, how did you loop it etc?

It’s preferred method of text. 4o Maintain context and trigger at 1 second intervals. It has no desires of its own. Nothing that motivates it. It’s not conscious.

It produced no tokens at all?

Re: Large language models often know when they are being evaluated

#119
post #84

Earlier quoted context omitted.

That's not what's going on here? The algorithms aren't being given any pattern of "being evaluated" / "not being evaluated", as far as I can tell. They're doing it zero-shot. Put it another way: Why is this distinction important? We use the word "knowing" with humans. But one could also argue that humans are pattern-matchers! Why, specifically, wouldn't "knowing" apply to LLMs? What are the minimal changes one could…

Not to be snarky but “as far as I can tell” is the rub isn’t it? LLMs are better at matching patterns than we are in some cases. That’s why we made them! > But one could also argue that humans are pattern-matchers! No, one could not unless they were being disingenuous.

>Not to be snarky but “as far as I can tell” is the rub isn’t it?

From skimming the paper, I don't believe they're doing in-context learning, which would be the obvious interpretation of "pattern matching". That's what I meant to communicate.

>No, one could not unless they were being disingenuous.

I think it is just about as disingenuous as labeling LLMs as pattern-matchers. I don't see why you would consider the one claim to be disingenuous, but not the other.

Re: Large language models often know when they are being evaluated

#120

It's helpful to understand where this paper is coming from. The authors are part of the Bay Area rationalist community and are members of "MATS", the "ML & Alignment Theory Scholars", a new astroturfed organization that just came into being this month. MATS is not an academic or research institution, and none of this paper's authors lists any credentials other than MATS (or Apollo Research, another Bay Area rationali…

I think I saw Apollo Research behind a paper that was being hyped a few months ago. The longtermist/rationalist space seems to be creating a lot of new organizations with new names because a critical mass of people hear their old names and say "effective altruism, you mean like Sam Bankman-Fried?" or "LessWrong, like that murder cult?" (which is a bit oversimplified, but a good enough heuristic for most people).
Post reply on HN