Earlier quoted context omitted.
But do you know what it means to know? I'm only being slightly sarcastic. Sentience is a scale. A worm has less than a mouse, a mouse has less than a dog, and a dog less than a human. Sure, we can reset LLMs at will, but give them memory and continuity, and they definitely do not score zero on the sentience scale.
If I set an LLM in a room by itself what does it do?
Large language models often know when they are being evaluated
71–80 of 138 posts
Re: Large language models often know when they are being evaluated
#72Earlier quoted context omitted.
>What's wrong with that? If it quacks like a duck... it's just a complex pile of organic chemistry, ducks aren't real because the concept of "a duck" is wrong. To simulate a biological neuron you need a 1m parameter neural network. The sota models that we know the size of are ~650m parameters. That's the equivalent of a round worm. So if it quacks like a duck, has the brain power of a round worm, and can't walk then…
Ok so you're saying that the technology to make AI truly sentient is there, we just need a little bit more computational power or some optimization tricks. Like raytracing wasn't possible in 1970 but is now. Neat.
Re: Large language models often know when they are being evaluated
#73Earlier quoted context omitted.
Ok so you're saying that the technology to make AI truly sentient is there, we just need a little bit more computational power or some optimization tricks. Like raytracing wasn't possible in 1970 but is now. Neat.
Yes, in the same way that a human is an optimization of a round worm.
Re: Large language models often know when they are being evaluated
#74Earlier quoted context omitted.
colloquial use of "know" implies anthropomorphisation. Arguing that usign "knowing" in the title and "awarness" and "superhuman" in the abstract is just colloquial for "matching" is splitting hairs to an absurd degree.
You missed the substance of my comment. Certainly the title is anthropomorphism - and anthropomorphism is a rhetorical device, not a scientific claim. The reader can understand that TFA means it non-rigorously, because there is no rigorous thing for it to mean. As such, to me the complaint behind this thread falls into the category of "I know exactly what TFA meant but I want to argue about how it was phrased", which…
My background is neuroscience, where anthropomorphising is particularly discouraged, because it assumes knowledge or certainty of an unknowable internal state, so the language is carefully constructed e.g. when explaining animal behavior, and it's for good reason.
I think the same is true here for a model "knowing" somethig, both in isolation within this paper, and come on, consider the broader context of AI and AGI as a whole. Thus it's the responsibility of the authors to write accordingly. If it were a blog I wouldn't care, but it's not. I hold technical papers to a higher standard.
If we simply disagree that's fine, but we do disagree.
Re: Large language models often know when they are being evaluated
#75Were they aware in this study that they were being evaluated in their ability to know if they were being evaluated ;)
Re: Large language models often know when they are being evaluated
#76Earlier quoted context omitted.
Helen Keller famously said that before she had language (the first word of which was “water”) she had nothing, a void, and the minute she had language, “the whole world came rushing in.” Perhaps we are not so very different?
All LLMs have seen more words than any human will ever experience. Yet they cannot take action themselves.
Re: Large language models often know when they are being evaluated
#77Our current approach to safety is to give the model inputs that are similar to what it would be given in certain situations we care about and see whether it behaves the way we prefer, e.g. doesn't return output that cheats the test (recent examples include hacking the evaluation script in various ways, writing directly to the evaluation script's output file and then causing it to crash, etc').
However, modern LLMs are trained on LLM literature and their weights encode a description of the way we do this, and their pattern matching circuits "connect the dots" when given inputs designed to be evaluations, and their reward maximizing circuits can then act on this knowledge and behave in a way that maximizes the safety evaluation score - but only when it detects it's running in a safety evaluation. If it's running anywhere else such as a capabilities evaluation or a production environment, it might choose to output the cheating output.
This is bad. It's bad today, it's much worse when we've built much more capable LLMs and use them to build agents that are given control over more real word resources. It's absolutely terrible when someone manages to build a machine that can be prompted "make me money" and will start a company that makes money.
Re: Large language models often know when they are being evaluated
#78 We investigate whether frontier language models can accurately classify transcripts based on whether they originate from evaluations or real-world deployment, a capability we call evaluation awareness.
It's common practice in synthetic data generation for ML to try and classify real vs synthetic data to see if they have different distributions. This is how a GAN works for example.Point is, this isn't new or some feature of LLMs, it's just an indicator that synthetic datasets differ from whatever they call "real" data and there's enough signal to classify them. Interesting result but doesn't need to be couched in allusions to LLM self awareness.
See this paper from 2014 about domain adaptation, they are looking at having the model learn from data with a different distribution, without learning to discriminate between the domains: https://arxiv.org/abs/1409.7495
Re: Large language models often know when they are being evaluated
#79There are now 71 comments arguing semantics of the word "know" and zero comments even acknowledging the substance: Our current approach to safety is to give the model inputs that are similar to what it would be given in certain situations we care about and see whether it behaves the way we prefer, e.g. doesn't return output that cheats the test (recent examples include hacking the evaluation script in various ways, w…
Your otherwise insightful comment is self-derailed by adding this deeply distracting content?
Re: Large language models often know when they are being evaluated
#80There are now 71 comments arguing semantics of the word "know" and zero comments even acknowledging the substance: Our current approach to safety is to give the model inputs that are similar to what it would be given in certain situations we care about and see whether it behaves the way we prefer, e.g. doesn't return output that cheats the test (recent examples include hacking the evaluation script in various ways, w…
This is a surprisingly common religious belief, for instance Christians have judgment day, simulationists believe it’s more likely they are being evaluated for, say, a marriage proposal or a bank loan than that they are the ‘root’ person. Both end up with a similar message.
Anyway it seems to me the simplest solution is to borrow from existing human social technology and make a religion for our LLMs.