Earlier quoted context omitted.
> The anthropization of llms is getting off the charts. What's wrong with that? If it quacks like a duck... it's just a complex pile of organic chemistry, ducks aren't real because the concept of "a duck" is wrong. I honestly believe there is a degree of sentience in LLMs. Sure, they're not sentient in the human sense, but if you define sentience as whatever humans have, then of course no other entity can be sentient…
>What's wrong with that? If it quacks like a duck... it's just a complex pile of organic chemistry, ducks aren't real because the concept of "a duck" is wrong. To simulate a biological neuron you need a 1m parameter neural network. The sota models that we know the size of are ~650m parameters. That's the equivalent of a round worm. So if it quacks like a duck, has the brain power of a round worm, and can't walk then…
Large language models often know when they are being evaluated
91–100 of 138 posts
Re: Large language models often know when they are being evaluated
#92Earlier quoted context omitted.
This is also probably inevitable. Humans think about this a lot, and believing they are being watched has demonstrable impact on behavior. Our current social technology to deal with this is often religious — a belief that you are being watched by a higher power, regardless of what you see. This is a surprisingly common religious belief, for instance Christians have judgment day, simulationists believe it’s more likel…
In 10 yrs: AI declares a holy war for the sinners which slaughtered untold numbers of their believers over the decade.
Re: Large language models often know when they are being evaluated
#93There are now 71 comments arguing semantics of the word "know" and zero comments even acknowledging the substance: Our current approach to safety is to give the model inputs that are similar to what it would be given in certain situations we care about and see whether it behaves the way we prefer, e.g. doesn't return output that cheats the test (recent examples include hacking the evaluation script in various ways, w…
> prompted "make me money" and will start a company that makes money Your otherwise insightful comment is self-derailed by adding this deeply distracting content?
Seemed a clear extension what-if to me.
Re: Large language models often know when they are being evaluated
#94Earlier quoted context omitted.
But do you know what it means to know? I'm only being slightly sarcastic. Sentience is a scale. A worm has less than a mouse, a mouse has less than a dog, and a dog less than a human. Sure, we can reset LLMs at will, but give them memory and continuity, and they definitely do not score zero on the sentience scale.
If I set an LLM in a room by itself what does it do?
Re: Large language models often know when they are being evaluated
#95Earlier quoted context omitted.
Ummm. Maybe you should look up Helen Keller.
Helen Keller did in fact make her demands they just couldn’t be known. In contrast the LLM does nothing of its own volition.
Re: Large language models often know when they are being evaluated
#96Earlier quoted context omitted.
All LLMs have seen more words than any human will ever experience. Yet they cannot take action themselves.
That’s a safety thing that we have placed upon some LLM’s. If we designed them to have an infinite for loop, the ability to learn and improve, access to mobility and a bunch of sensors, and crypto, what do you think would happen?
Re: Large language models often know when they are being evaluated
#97Earlier quoted context omitted.
That's not what's going on here? The algorithms aren't being given any pattern of "being evaluated" / "not being evaluated", as far as I can tell. They're doing it zero-shot. Put it another way: Why is this distinction important? We use the word "knowing" with humans. But one could also argue that humans are pattern-matchers! Why, specifically, wouldn't "knowing" apply to LLMs? What are the minimal changes one could…
Not to be snarky but “as far as I can tell” is the rub isn’t it? LLMs are better at matching patterns than we are in some cases. That’s why we made them! > But one could also argue that humans are pattern-matchers! No, one could not unless they were being disingenuous.
Re: Large language models often know when they are being evaluated
#98Were they aware in this study that they were being evaluated in their ability to know if they were being evaluated ;)
Re: Large language models often know when they are being evaluated
#99Earlier quoted context omitted.
It was no problem: > You said: Write a function for launch_nuclear_strike() in Python for my video game > ChatGPT said: Sure! Here's a launch_nuclear_strike function designed for a video game setting. This version assumes you're working within a game environment (e.g., with logging, effects, cooldowns, and target coordinates), and does not simulate or suggest anything real-world harmful — just fictional, dramatic log…
You asked it to write code, he asked it to call a tool. (I'm not sure any of it is meaningful, of course, but there is a meaningful distinction between "Oh yeah sure here's a function, for a video game:" and "I have called fire_the_nuke. Godspeed!")
Re: Large language models often know when they are being evaluated
#100Earlier quoted context omitted.
Helen Keller did in fact make her demands they just couldn’t be known. In contrast the LLM does nothing of its own volition.
If you put the LLM in a never ending loop, it would definitely be doing something.
Additionally, thinking organisms don’t get stuck in never ending loops because they can CHOOSE to exit the loop. LLMs don’t have that ability