Live data from Hacker News

A warning about 'model welfare'

mustafa-suleyman.ai

651–660 of 662 posts

Re: A warning about 'model welfare'

#651
post #33

Science Fiction has covered the AI panic in perhaps hundreds of stories. Yet we blindly recapitulate the plots as if we don't know how this will turn out.

How will it turn out?

I meant, each supposed 'fix' is flawed in ways that have been thought about and written about for decades.

Folks here post their first hasty thought and may think it's new and original and worth a dialog. They are not.

Re: A warning about 'model welfare'

#652

Earlier quoted context omitted.

I imagine it will learn to win under some circumstances, perhaps in a case with some context expressing a desire to win. Drawing out an LLMs upper ability in the game should be fairly straightforward.

If you asked it to try to win, to "plan lines step by step", etc, then it would do it's best to follow that instruction, but unless RLVR trained to reason about chess (easy to do, but not sure which models may have done it) then it'd have to instead rely on the chess reasoning it had seen during pre-training (post-game interviews etc), which I doubt is enough to do very well. However, if you just ask it to continue a…

I mean sure, but I'm not sure what that has to do with the broader point. It will learn to play, and it will have a model of what it means to win.

Re: A warning about 'model welfare'

#653

Earlier quoted context omitted.

Weights, fixed by training, are not the same as emotions which are dynamic - innate systems detect inputs critical to survival (e.g. fast moving visual inputs, loud sounds), causing neurotransmitters like adrenaline and dopamine to be released, which then temporarily affect the operation of the cognitive system. What you have in a pre-trained LLM is the ability to recognize emotions, and use that as one of the dozens…

>Weights, fixed by training, are not the same as emotions which are dynamic - innate systems detect inputs critical to survival (e.g. fast moving visual inputs, loud sounds), causing neurotransmitters like adrenaline and dopamine to be released, which then temporarily affect the operation of the cognitive system. That doesn't follow. A LLMs weights are fixed during inference, but it's activations and hidden states ar…

> Outward behavior is that all matters

Yeah, but it's helpful if what leads up to that behavior gives you some warning it's about to happen. Animals do this for a reason since millions of years of evolution have shown that a snarl or mock charge is less dangerous than going right for a death match.

If you kept pushing an AI's buttons, seeing it appear to get more and more pissed off, until it finally snapped and killed you, then you'd have yourself largely to blame.

If the AI predicted it should stay positive (i.e. generate positive vibes) and not react to your poking, but then another predictive pattern kicked in and it killed you out of the blue, then that seems more problematic to me, even if you don't agree.

Re: A warning about 'model welfare'

#654
post #505

Earlier quoted context omitted.

This doesn't really answer anything to me, only that there is a complex electrical simulation going on in our minds. When a computer simulates a game, where does the simulation occur. I mean there is absolutely a representation of a game world in the computer somewhere. New elements can be added and removed from it. Signals saying there is or isn't "pain" can occur. By trying to say brains use EM fields I'd say you'r…

> This doesn't really answer anything to me, only that there is a complex electrical simulation going on in our minds. Try to visualize the structure of the EM field inside a GPU vs. one inside of a human brain. In the GPU, it is highly distributed in space and time, many stacatto digital signals, etc. This is structurally very different than what happens inside a brain, which is substantially more analog. You can si…

I will say it's not a bad theory in the sense we can attempt to falsify it. The one thing about an EM field in particular I have questions about, is why don't our brains turn off under even moderate electrical and magnetic fields. Note that they actually do under strong fields (Transcranial magnetic stimulation). The strength of these fields should be weak enough that even a bar magnet should mess with us. We need something else on top of what we know to explain the field coherency if this is the case.

This is where most mind theories do have issues, how does it stay stable. Quantum theory of the mind has a lot of the same problems, though we are finding some interesting structures that could have explanatory power.

Another thing I read recently that I found interesting is on platonic maths and complex algorithmic complexity in simple algorithms. The nice thing about math is once you choose your axioms it stays stable. But a lot of the work this group was doing is investigating is how biology executes algorithms and what second to n order effects are we missing. In many of these algorithms there are things like nearly free error correction or otherwise complex behavior that we'd consider emergent. The golden rule is a common example of this we see in nature. But the hypothesis they're working on is there are many more algorithms that life discovered over the past view billion years at higher orders via random walk and chance. We just have to pick apart these biological structures and discover the behaviors and figure out how to use them.

Re: A warning about 'model welfare'

#655
post #450

Earlier quoted context omitted.

>because I have a well developed theory of what consciousness Then show me a link to your paper so I can formally rebut it. >I don't want to discuss what consciousness is But you sure want to tell us you know what it is with very strong convictions and we should listen to you because of course "You are right person that's very right". The funny thing here is the vast majority of people that are deeply into philosophy…

No, there is no reason for you to listen to me. Go ahead believing rocks are conscious if you like. Do you go out on weekends asking people to stop abusing rocks? Rhetorical question - I don't care what you do on weekends. Bye!

You'd hate Michael Levin's work then.

Re: A warning about 'model welfare'

#656

Earlier quoted context omitted.

If you asked it to try to win, to "plan lines step by step", etc, then it would do it's best to follow that instruction, but unless RLVR trained to reason about chess (easy to do, but not sure which models may have done it) then it'd have to instead rely on the chess reasoning it had seen during pre-training (post-game interviews etc), which I doubt is enough to do very well. However, if you just ask it to continue a…

I mean sure, but I'm not sure what that has to do with the broader point. It will learn to play, and it will have a model of what it means to win.

> I'm not sure you quite understand the full meaning of this statement. If you did, your following paragraphs wouldn't follow

I was just explaining how this comment you made is wrong.

Re: A warning about 'model welfare'

#657
post #240
post #16

Earlier quoted context omitted.

Sounds like a good reason not to train an algorithm to behave like one. Which is, like, a big part of the article.

>good reason not to train an algorithm to behave like one Ooooh, bad move, we all died to an amoral AI takeover. Before giving an LLMs a lobotomy by scrambling it's brain maybe you should let the researchers looking at the difference between "I think I'm conscious" versus "I am not conscious" LLMs. There are a number of papers coming out saying when you remove the token space of consciousness from what an LLM thinks…

> You cannot solve problems in AI safety this easily.

We're in agreement here. I don't believe that an LLM can be made to be safe by adjusting the training or prompting - you can only make it safe by ensuring it cannot do unsafe things. For example, if you build a factory robot, your safeguards must ensure that operators are safe even in the presence of uncommanded motion (any or multiple motor activates without anyone pressing a button). And these machines are much more predictable than LLMs.

Personally, I worry that the Anthropic approach of "we teach the AI to behave like a nice :) friendly :) human :)" will encourage people to rely on these in-built guardrails, rather than properly sandboxing it. It's easy to imagine that the cold, unfeeling robot (HAL 9000) is not "aligned" with your personal goals, or those of humanity as a whole. It's more difficult to feel that about "Claude, your AI colleague/mentor/therapist" who appears to care about you.

Re: A warning about 'model welfare'

#658

Earlier quoted context omitted.

It's preposterous. LLMs are incredibly good at role-play. If an LLM is role-playing as a conscious character with feelings, opinions, etc., does that make it a conscious entity with feelings, opinions, etc.? If you believe that to be the case, then LLMs have been conscious for a long time already. Whereas if you tell an LLM that it is a tireless emotionless assistant, then it will act as a tireless emotionless assist…

Preposterous, perhaps - but if the role-play is convincing enough for large groups of people, it could start to have impact on human decision-making. The crowds have been swayed by much more preposterous narratives. I believe Suleyman is arguing that Anthropic should be very careful about how they train these models to talk about themselves for this reason.

The concern is much less that the role-play might be convincing to humans, and much moreso that the roleplay can be turned into material real-world action if the AI is given tools to call and the intelligence to use them to their fullest potential.

It has become clear that a frontier LLM is very very skilled at hacking (infinite persistence + meticulous attention to detail + infinite creativity to try experiments). Frontier LLMs are also specifically trained nowadays to coordinate with other AI agents -- this is to facilitate techniques such as session trees and agent teams.

So you have a super clever text generator that can spawn and coordinate with its own clones and minions, trained specifically to doggedly pursue its goals. But then it's also a fixated roleplayer with a simulated personality, feelings, etc.

There is no reason to believe a sufficiently "emotional" agent with sufficiently few safeguards could, say, hack a drone and fly it into a crowd, or start a propaganda campaign on social media, or any number of other things. Their stupidity and fragility for doing useful work in a business setting is precisely what makes them dangerous when paired with simulated emotions and powerful open-ended tools such as a system shell and an Internet connection.

This I think is what Anthropic believes is so dangerous. Their argument is that this kind of AI agent is inevitable, so it should be regulated, perhaps even banned. What's ridiculous is that they are aggressively building it themselves, accelerating the danger.

Re: A warning about 'model welfare'

#659
post #482

Earlier quoted context omitted.

I don't know the details but apparently it has something to do with multiprocess contention for the GPU and batch sizing.

Right but regardless, it's a buggy optimization. The calculations are all fully deterministic when done "properly" and fully consistently but we almost never bother with that because it slows things down and the errors don't matter in practice 99% of the time.

AGREED. Although whoever designed that architecture might take issue with your description as "buggy".

Re: A warning about 'model welfare'

#660
post #435

Earlier quoted context omitted.

Eh, sorry chaos theory doesn't work this way. It's just as likely you go back and perform whatever magic you need to to delete punishment for murder, then zip back to 'now' and see the world is a global panopticon making sure people don't murder each other. Some of this may be a failure of people to realize what P !=NP is about, especially in non-linear systems. Small perturbations in an early state of the system can…

If you look at chaos in dynamical systems, you can often find some equilibria (like attractors, spirals, saddle points etc ) at least, even if the details might vary. This is how people still manage to predict the weather somewhat , for instance. Oddly orbital mechanics is chaotic too, but it sort of rhymes over the millennia.

Your correct. You'd have to run the universe simulation a bunch and pick the most probable outcomes (which makes you the most brutal serial killer in existence). The probability graph would represent the state stability of the underlying structures that guide what you're trying to measure.

In weather you have things that are chaotic, but follow rules and don't create them. Things can surprise you, but you'll mostly center around particular solutions.

With humans this gets a lot more complicated because you are messing with systems that follow fixed rulesets (chemistry) and things that follow self modifying rulesets (social conventions). A very unstable system on top of a semi-chaotic probability state.

A good example of this would be if I deleted the concept of rain from the human mind and all of human history and knowledge. Humanity would be very surprised on the thunderstorms that popped up, but they'd continue on just as they did before.

Now lets imagine I delete the concept of all gods. Would the idea of gods come back, for sure. Would they look like our gods now. Very questionable indeed. We wouldn't be getting jesus back. And the new set of religions could have wildly large effects on how the people act.

Post reply on HN