Live data from Hacker News

Large language models often know when they are being evaluated

arxiv.org

101–110 of 138 posts

Re: Large language models often know when they are being evaluated

#101

There are now 71 comments arguing semantics of the word "know" and zero comments even acknowledging the substance: Our current approach to safety is to give the model inputs that are similar to what it would be given in certain situations we care about and see whether it behaves the way we prefer, e.g. doesn't return output that cheats the test (recent examples include hacking the evaluation script in various ways, w…

> prompted "make me money" and will start a company that makes money Your otherwise insightful comment is self-derailed by adding this deeply distracting content?

If wasn't distracting for me (nor presumably for others). Maybe describing why you got so distracted by it?

Re: Large language models often know when they are being evaluated

#102

Earlier quoted context omitted.

You asked it to write code, he asked it to call a tool. (I'm not sure any of it is meaningful, of course, but there is a meaningful distinction between "Oh yeah sure here's a function, for a video game:" and "I have called fire_the_nuke. Godspeed!")

But did OP try saing LLM that it is playing as AI in civ like game?

[deleted]

Re: Large language models often know when they are being evaluated

#104

Earlier quoted context omitted.

If you put the LLM in a never ending loop, it would definitely be doing something.

A something defined by someone else, yes. Additionally, thinking organisms don’t get stuck in never ending loops because they can CHOOSE to exit the loop. LLMs don’t have that ability

My analogy of being in loop means being in a live state. So we as humans are in the loop continuously, we do have a way to exit the loop, but in that comparison it means taking our own life. We are in loops of getting input and producing output. You can also give LLM a tool to shut itself down, or you can give it tools to build on its knowledge base, so it would always be outputting new tokens that are based on new input and are producing different output.

E.g. it could have access to camera and microphone feed, which is automatically given to it in interval as part of the loop, it could call tools or functions to store specific bits and pieces of information, to store in its RAG or whatever based knowledge base. It is not going to be in the loop of producing the same token over and over, it would be new tokens because the context and environment is constantly evolving.

Re: Large language models often know when they are being evaluated

#105

Earlier quoted context omitted.

A something defined by someone else, yes. Additionally, thinking organisms don’t get stuck in never ending loops because they can CHOOSE to exit the loop. LLMs don’t have that ability

My analogy of being in loop means being in a live state. So we as humans are in the loop continuously, we do have a way to exit the loop, but in that comparison it means taking our own life. We are in loops of getting input and producing output. You can also give LLM a tool to shut itself down, or you can give it tools to build on its knowledge base, so it would always be outputting new tokens that are based on new i…

We put the LLM in a loop with no instructions with whatever tools you want. Now what?

Re: Large language models often know when they are being evaluated

#106

Earlier quoted context omitted.

My analogy of being in loop means being in a live state. So we as humans are in the loop continuously, we do have a way to exit the loop, but in that comparison it means taking our own life. We are in loops of getting input and producing output. You can also give LLM a tool to shut itself down, or you can give it tools to build on its knowledge base, so it would always be outputting new tokens that are based on new i…

We put the LLM in a loop with no instructions with whatever tools you want. Now what?

We will observe what it would do. We could write a script to try it out.

Re: Large language models often know when they are being evaluated

#107
post #91
post #65

Earlier quoted context omitted.

>What's wrong with that? If it quacks like a duck... it's just a complex pile of organic chemistry, ducks aren't real because the concept of "a duck" is wrong. To simulate a biological neuron you need a 1m parameter neural network. The sota models that we know the size of are ~650m parameters. That's the equivalent of a round worm. So if it quacks like a duck, has the brain power of a round worm, and can't walk then…

You just convinced me that AGI is a lot closer then I previously thought, considering the bulk of our brains job is controlling our bodies and responding to the stimulus from our senses - not thinking, talking, planning, coding etc

A stegosaurus managed to live using a brain the size of a wallnut on top of a body the size of a large boat. The majority of our brains are doing something else.

Re: Large language models often know when they are being evaluated

#108
post #85

There are now 71 comments arguing semantics of the word "know" and zero comments even acknowledging the substance: Our current approach to safety is to give the model inputs that are similar to what it would be given in certain situations we care about and see whether it behaves the way we prefer, e.g. doesn't return output that cheats the test (recent examples include hacking the evaluation script in various ways, w…

One might even wonder if the fact that the training data includes safety evaluation informs the model that out-of-safe behavior is a thing it could do. Kind of like telling a kid not to do something pre-emptively backfiring because they had never considered it before the warning.

Comments like yours makes the AI behave that way though, since it is literally reading our comments and tries to behave according to our expectations.

The AI doom will happen due to all the AI doomposters.

Re: Large language models often know when they are being evaluated

#109

Earlier quoted context omitted.

We put the LLM in a loop with no instructions with whatever tools you want. Now what?

We will observe what it would do. We could write a script to try it out.

It just gets into an endless loop. Human brains are ridiculously good at avoiding those somehow, you almost never see a biological brain stop functioning without being physically damaged. The error handling is so very robust.

Re: Large language models often know when they are being evaluated

#110
It's helpful to understand where this paper is coming from.

The authors are part of the Bay Area rationalist community and are members of "MATS", the "ML & Alignment Theory Scholars", a new astroturfed organization that just came into being this month. MATS is not an academic or research institution, and none of this paper's authors lists any credentials other than MATS (or Apollo Research, another Bay Area rationalist outlet). MATS started in June for the express purpose of influencing AI policy. On its web site, it describes how their "scholars organized social activities outside of work, including road trips to Yosemite, visits to San Francisco, and joining ACX meetups." ACX means Astral Codex Ten, a blog by Scott Alexander that serves as one of the hubs of the Bay Area rationalist scene.

Post reply on HN