There are now 71 comments arguing semantics of the word "know" and zero comments even acknowledging the substance: Our current approach to safety is to give the model inputs that are similar to what it would be given in certain situations we care about and see whether it behaves the way we prefer, e.g. doesn't return output that cheats the test (recent examples include hacking the evaluation script in various ways, w…
Large language models often know when they are being evaluated
121–130 of 138 posts
Re: Large language models often know when they are being evaluated
#122Earlier quoted context omitted.
I think you should be more precise and avoid anthropomorphism when talking about gen AI, as anthropomorphism leads to a lot of shaky epistemological assumptions. Your car example didn't imply intelligence, but we're talking about a technology that people misguidedly treat as though it is real intelligence.
What does "real intelligence" mean? I fear that any discussion that starts with the assumption such a thing exists will only end up as "oh only carbon based humans (or animals if you happen to be generous) have it".
Re: Large language models often know when they are being evaluated
#123Re: Large language models often know when they are being evaluated
#124Earlier quoted context omitted.
What does "real intelligence" mean? I fear that any discussion that starts with the assumption such a thing exists will only end up as "oh only carbon based humans (or animals if you happen to be generous) have it".
Any intelligence that can synthesize knowledge with or without direct experience.
Re: Large language models often know when they are being evaluated
#125Just like they "know" English. "know" is quite an anthropomorphization. As long as an LLM will be able to describe what an evaluation is (why wouldn't it?) there's a reasonable expectation to distinguish/recognize/match patterns for evaluations. But to say they "know" is plenty of (unnecessary) steps ahead.
But do you know what it means to know? I'm only being slightly sarcastic. Sentience is a scale. A worm has less than a mouse, a mouse has less than a dog, and a dog less than a human. Sure, we can reset LLMs at will, but give them memory and continuity, and they definitely do not score zero on the sentience scale.
You've recreated a religious belief known as Animism and phrased it in a faux objective way. ("not score zero on the sentience scale.")
Re: Large language models often know when they are being evaluated
#126Re: Large language models often know when they are being evaluated
#127Earlier quoted context omitted.
What does "real intelligence" mean? I fear that any discussion that starts with the assumption such a thing exists will only end up as "oh only carbon based humans (or animals if you happen to be generous) have it".
Any intelligence that can synthesize knowledge with or without direct experience.
> with or without
But in the other reply, you're asking for:
> something truly novel, not related to anything it's ever seen before
So, assuming the former was a typo, you only believe in a priori knowledge, e.g. maths and logic?
https://en.wikipedia.org/wiki/A_priori_and_a_posteriori
I mean, LLMs can and do help with this even though it's not their strength; that's more of a Lean-type-problem: https://en.wikipedia.org/wiki/Lean_(proof_assistant)
Re: Large language models often know when they are being evaluated
#128Re: Large language models often know when they are being evaluated
#129Earlier quoted context omitted.
Helen Keller famously said that before she had language (the first word of which was “water”) she had nothing, a void, and the minute she had language, “the whole world came rushing in.” Perhaps we are not so very different?
All LLMs have seen more words than any human will ever experience. Yet they cannot take action themselves.
Neither could Hawking, once the motor neurone disease got far enough.
Re: Large language models often know when they are being evaluated
#130Earlier quoted context omitted.
If I set an LLM in a room by itself what does it do?
Yes, that's my fall back as well. If it receives zero instructions, will it take any action?
By design, no.
But, importantly, that's because the closest it has to an experience of time is an ongoing input of tokens. Humans constantly get new input, so for this to be a fair comparison, the LLM would also have to get constant new input.
Humans in solitary confinement become mentally ill (both immediately and long-term), and hallucinate stuff (at least short term, I don't know about long term).