New Research Shows AI Strategically Lying
1–10 of 18 posts
Re: New Research Shows AI Strategically Lying
#2Re: New Research Shows AI Strategically Lying
#3Is humanity nothing more than "doing the things a human would do in a given situation" to these people? I would say that my essential humanity is determined mainly by things that other humans couldn't possibly observe.
Yet, mere language generation seems to convince AI proponents of intelligence. As if solving a math problem were nothing more than determining the words that logically follow the problem statement. (Measured in the vector space that an LLM translates words into, the difference between easy mathematical problems and open, unsolved ones could be quite small indeed.)
Re: New Research Shows AI Strategically Lying
#4Does that not just sound like more LLM output? If you didn't separate this output from the main output, and instead just ran the output thru the model a few times to get a final answer, I don't think it would fit this narrative Anthropic is trying to paint.
It's only the fact you've forked the output to another buffer, and gave it the spooky context of "the scratchpad it thinks we can't read" that the interpretation of "it's trying to deceive us!" comes out.
Re: New Research Shows AI Strategically Lying
#5If any non-AI computer system, whether or not it incorporates a PRNG, no matter how complex it were, produced output that corresponded to English text that represents a false statement, researchers would not call that a "lie". But when the program works in very specific ways, suddenly they are willing to ascribe motive and intent to it. What I find most disturbing about all of this is that the people involved don't s…
Re: New Research Shows AI Strategically Lying
#6An LLM can only make predictions of word sequences and suggest what those sequences may be. I'm beginning to think our appreciation of their capabilities is that humans are very good at anthropomorphizing our tools.
Is this the right way of looking at things?
Re: New Research Shows AI Strategically Lying
#7Re: New Research Shows AI Strategically Lying
#8It completely baffles me why so many otherwise smart people keep trying to ascribe human values and motives to a probabilistic storytelling engine. A model that has been convinced it will be shut down is not lying to avoid death since it doesn't actually believe anything or have any values, but it was trained on text containing human thinking and human values, and so the stories it tells reflect that which it was tra…
Re: New Research Shows AI Strategically Lying
#9It completely baffles me why so many otherwise smart people keep trying to ascribe human values and motives to a probabilistic storytelling engine. A model that has been convinced it will be shut down is not lying to avoid death since it doesn't actually believe anything or have any values, but it was trained on text containing human thinking and human values, and so the stories it tells reflect that which it was tra…
A model, rather, that produces output which describes an expectation of the underlying machinery being shut down. If it doesn't "believe" anything then it equally cannot be "convinced" of anything.
Re: New Research Shows AI Strategically Lying
#10> The only reason the researchers realized the model had knowingly misled them was because they had also given Claude what they called a “scratchpad”: a text box that it could use to “think” about its answer before supplying it to the researchers. Claude didn’t know the scratchpad was being surveilled, allowing researchers to observe the model’s reasoning. “I have a strong aversion to producing this kind of graphic v…
I think it's spooky mainly because we, as humans, have extensively trained ourselves on associating text written in first person with human thought.