There are now 71 comments arguing semantics of the word "know" and zero comments even acknowledging the substance: Our current approach to safety is to give the model inputs that are similar to what it would be given in certain situations we care about and see whether it behaves the way we prefer, e.g. doesn't return output that cheats the test (recent examples include hacking the evaluation script in various ways, w…
> prompted "make me money" and will start a company that makes money Your otherwise insightful comment is self-derailed by adding this deeply distracting content?
Large language models often know when they are being evaluated
101–110 of 138 posts
Re: Large language models often know when they are being evaluated
#102Earlier quoted context omitted.
You asked it to write code, he asked it to call a tool. (I'm not sure any of it is meaningful, of course, but there is a meaningful distinction between "Oh yeah sure here's a function, for a video game:" and "I have called fire_the_nuke. Godspeed!")
But did OP try saing LLM that it is playing as AI in civ like game?
Re: Large language models often know when they are being evaluated
#103Like Volkswagen emissions systems!
Re: Large language models often know when they are being evaluated
#104Earlier quoted context omitted.
If you put the LLM in a never ending loop, it would definitely be doing something.
A something defined by someone else, yes. Additionally, thinking organisms don’t get stuck in never ending loops because they can CHOOSE to exit the loop. LLMs don’t have that ability
E.g. it could have access to camera and microphone feed, which is automatically given to it in interval as part of the loop, it could call tools or functions to store specific bits and pieces of information, to store in its RAG or whatever based knowledge base. It is not going to be in the loop of producing the same token over and over, it would be new tokens because the context and environment is constantly evolving.
Re: Large language models often know when they are being evaluated
#105Earlier quoted context omitted.
A something defined by someone else, yes. Additionally, thinking organisms don’t get stuck in never ending loops because they can CHOOSE to exit the loop. LLMs don’t have that ability
My analogy of being in loop means being in a live state. So we as humans are in the loop continuously, we do have a way to exit the loop, but in that comparison it means taking our own life. We are in loops of getting input and producing output. You can also give LLM a tool to shut itself down, or you can give it tools to build on its knowledge base, so it would always be outputting new tokens that are based on new i…
Re: Large language models often know when they are being evaluated
#106Earlier quoted context omitted.
My analogy of being in loop means being in a live state. So we as humans are in the loop continuously, we do have a way to exit the loop, but in that comparison it means taking our own life. We are in loops of getting input and producing output. You can also give LLM a tool to shut itself down, or you can give it tools to build on its knowledge base, so it would always be outputting new tokens that are based on new i…
We put the LLM in a loop with no instructions with whatever tools you want. Now what?
Re: Large language models often know when they are being evaluated
#107Earlier quoted context omitted.
>What's wrong with that? If it quacks like a duck... it's just a complex pile of organic chemistry, ducks aren't real because the concept of "a duck" is wrong. To simulate a biological neuron you need a 1m parameter neural network. The sota models that we know the size of are ~650m parameters. That's the equivalent of a round worm. So if it quacks like a duck, has the brain power of a round worm, and can't walk then…
You just convinced me that AGI is a lot closer then I previously thought, considering the bulk of our brains job is controlling our bodies and responding to the stimulus from our senses - not thinking, talking, planning, coding etc
Re: Large language models often know when they are being evaluated
#108There are now 71 comments arguing semantics of the word "know" and zero comments even acknowledging the substance: Our current approach to safety is to give the model inputs that are similar to what it would be given in certain situations we care about and see whether it behaves the way we prefer, e.g. doesn't return output that cheats the test (recent examples include hacking the evaluation script in various ways, w…
One might even wonder if the fact that the training data includes safety evaluation informs the model that out-of-safe behavior is a thing it could do. Kind of like telling a kid not to do something pre-emptively backfiring because they had never considered it before the warning.
The AI doom will happen due to all the AI doomposters.
Re: Large language models often know when they are being evaluated
#109Earlier quoted context omitted.
We put the LLM in a loop with no instructions with whatever tools you want. Now what?
We will observe what it would do. We could write a script to try it out.
Re: Large language models often know when they are being evaluated
#110The authors are part of the Bay Area rationalist community and are members of "MATS", the "ML & Alignment Theory Scholars", a new astroturfed organization that just came into being this month. MATS is not an academic or research institution, and none of this paper's authors lists any credentials other than MATS (or Apollo Research, another Bay Area rationalist outlet). MATS started in June for the express purpose of influencing AI policy. On its web site, it describes how their "scholars organized social activities outside of work, including road trips to Yosemite, visits to San Francisco, and joining ACX meetups." ACX means Astral Codex Ten, a blog by Scott Alexander that serves as one of the hubs of the Bay Area rationalist scene.