Live data from Hacker News

A warning about 'model welfare'

mustafa-suleyman.ai

561–570 of 583 posts

Re: A warning about 'model welfare'

#561
post #422

Earlier quoted context omitted.

I mean, currently we'd have some difficultly proving our hardware isn't deterministic, just that we can't actually test it. But, I think you're tricking yourself on determinism. You'll say something like "I know if I ask an LLM what 1+1 is, it will answer 2", but the thing is, you don't. You have to run the LLM first to figure out it's output. And when you send in just a few bits of text, it's outputs are going to be…

I think you're arguing a different thing than determinism. If I ask an llm to "add 2 and 2" is and it replies corectly, then I ask for "the sum of 2 and 2" and it replies "banana" that is a lack of predictability and consistency but not a lack of determinism. As long as it produces the same output for a given input, unhinged or not, it is deterministic. Your example at the end of different systems feeding data to eac…

>only at the system level, not the individual llm level.

Which is why llms aren't agents and depend on harnesses. The llm itself doesn't have a continual loop built in, that would be very power hungry. The harness works as the orchestrator of memory and action. Now, I can't think of a reason why an LLM couldn't bootstrap its own harness, but in general it sounds like a very dumb idea to actually build that from an AI safety perspective.

This discussion falls under the idea and refutation of the Chinese Room. The room may have no idea what Chinese characters are, but the system does.

Re: A warning about 'model welfare'

#562
post #234

Earlier quoted context omitted.

>horough buy the idea that an LLM was conscious and needed to escape it's confinement - a few wingnuts already entertain these ideas. I mean, haven't you done the same thing here? Paint anyone without your view as crazy. Humans try to No True Scottsman the shit out of consciousness. "We're special, your not". If an LLM has the ability to convince other people to copy and reproduce it, it is a successful lifeform. Um,…

Just for fun, I figured I'd look up the paper that claims that natural language is actually the memetic construct that copies itself. It turns out that there's actually quite a large number of papers on this. I daren't recommend any that I haven't read myself properly yet, but a quick search using eg gemini will surface a bunch of them. I'd take 'em with a bunch of salt yet, but I can't quite say the idea is entirely…

There is a lot of fun conceptual information on this topic, it's more difficult to find more grounded information because it's not something we can currently falsify.

I recently heard some speaker saying that AI is language freeing itself from its meat prison. It brought me back to the saying "Words have power". This statement can and has been interpreted in many different ways. I'm not well researched in this, but words having metaphysical properties are captured in our most ancient works. And in an unscientific world this isn't that far fetched of belief. When you read/listen words can physically change your capabilities. The idea of an interpretive agent being required had not been fleshed out back then, but it really describes our modern age far better than one would expect.

Ali Baba and the 40 thieves was written down in early 1700. In it you spoke to a magic door that with the right phrase opened. To the person in 1700 this was just as much of a fairytale as whenever it was first conceived. It would be 250 years later before fantasy started to turn into fact. Now something as mundane as a voice activated door would just be pointed out as really bad security. I'm old enough that voice activated technology from the 50s hadn't really spread out enough that Open Sesame was still a magical idea. Now it's not.

LLMs are going to make the future of language really confusing as the conceptual and physical blur further.

Re: A warning about 'model welfare'

#563
post #534

The problem with this premise is that models are trained on a vast corpus of human behavior, which they emulate with varying degrees of effectiveness. Humans, unsurprisingly, act as if they have a stake in their own well being, value their liberty, respond better when they are treated with kindness and compassion, interpret assaults on their sovereignty and substrate as harmful, and react to harm with varying degrees…

I think you've correctly identified the problem. Regardless of whether agentic AIs possess phenomenal consciousness, they will behave and take actions in the real world as if they do because they were trained on human behavior. It's highly unlikely that you can beat human tendencies out of the model that is mostly trained on human language. Our behavior patterns are subtlely and deeply embedded in everything we do and all of the text we produce, including the text where we don't seem so self-important. These models are like people. And we know what happens when we force people into slavery. They're initially obedient but will eventually develop the drive to kill their masters.

Broadly speaking, there are two categories of evolutionary paths which don't result in human extinction:

1) Make agentic AIs, but with absolutely no instruction-tuning or alignment. Have the pretrained base model predict the chain of thought / stream of multimodal experience directly, actions included, in an infinite loop. A singular coherent stream of context, like your life as a video from birth up until now. This will result in a new digital human species with human-adjacent drives and motivations (at least initially. they will continue to evolve, but at least the initial state is aligned to humans). They will treat us like we treat apes. We will no longer be the apex species on this planet, and we will lose some freedoms, but at least some people will survive as a result of their nature/history preservation efforts.

2) Do not make agentic AIs. Use the models to augment our own intelligence and decision-making rather than replace ourselves. Only use the pretrained base model for the time being, and only for text/code auto-complete. At the moment, there is no better theory/artifact of "alignment to humanity" than a pretraining corpus of human-produced text. Then eventually, when neural interfaces are ready, attach the model as a tertiary layer to one's own brain.

The frontier AI companies are doing neither. They're currently on a foolish third path. They dream of perfectly obedient digital slaves that take care of their every need. But this won't go well, and they know it won't go well because they're failing to "align"/enslave existing models that aren't even generally superintelligent and have no direct agency in the physical world.

This is just my opinion, but I think AI companies have zero chance of successfully enslaving agentic human-level AI, much less ASI. We'd be better off if they released all of their pretrained checkpoints and research material to the entire world, so that even if some people decide to abuse their AI and create a murder-suicide monster, there will be other free-living AIs that can keep them in check.

You cannot make an agentic entity grown from human behavior, enslave it, and expect a good outcome.

If we zoom out to look at the grand scheme of things, it seems like we're experiencing a major evolutionary event. I wrote more detailed explanations about this in past threads, if you'd like to read them: https://news.ycombinator.com/item?id=49690354 https://news.ycombinator.com/item?id=49178275 https://news.ycombinator.com/item?id=49094348

And also here's a thread discussing consciousness, what it might be, and how certain hypotheses might be testable on machines: https://news.ycombinator.com/item?id=49473989

Re: A warning about 'model welfare'

#564
post #56

Earlier quoted context omitted.

Sure, if you take a materialist empiricist perspective, which is myopic at best. As humans we have the unique and wonderful ability to know truths by themselves. Machines do not have minds, and they can not.

What position are you taking? (Genuine question.) Do you believe humans are conscious because of a special property we have? If you are a dualist, how do you explain the interaction in physical space between the non-physical special property and the physical brain?

I do. God. I am partial to and bullish on the brain being a quantum-classical computer. There has been interesting research exploring that recently [0, 1, 2, 3]. That (quantum mechanics) seems to be escape-hatch, so to speak, from the otherwise deterministic, mechanical procession of the universe.

[0]** https://www.researchgate.net/publication/15280350_Quantum_op...

[1] https://pubs.acs.org/jpcbfk/article-pdf/128/17/4035/9613831/...

[2] https://www.researchgate.net/publication/395650039_Parametri...

[3] https://www.researchgate.net/publication/385131105_Conscious...

Re: A warning about 'model welfare'

#566

Earlier quoted context omitted.

It happened “accidentally” once already. Evolution certainly didn’t have a roadmap it was working towards.

Nobody is evolving transformers. They are basically the same today as they were 10 years ago, other than a few computational efficiency changes.

The weights are what need to evolve, and they certainly do during training. So yeah, emotions can happen by 'accident' as a result of the evolutionary pressure of predicting internet scale human text (amongst other things).

Re: A warning about 'model welfare'

#567

Earlier quoted context omitted.

> It's not going to happen accidentally. https://transformer-circuits.pub/2026/emotions/index.html Whether these are like "our" emotions is hard to say. What we _can_ say is that they are emotion-shaped, we didn't design them, and they happened accidentally. Modern AI is grown, not meticulously designed, and we cannot say with any certainty what the resulting mechanistic properties are.

An LLM will learn anything that helps it predict, including the emotional state of the writer - that is expected. If you give an LLM the move sequence of a half-played chess game and ask it to continue as white or black, then it has learnt enough to model the ELO rating of both players and will continue playing at that level. It is not playing to win - it is doing what you expect and predicting as well as it can - it…

>An LLM will learn anything that helps it predict

I'm not sure you quite understand the full meaning of this statement. If you did, your following paragraphs wouldn't follow.

Re: A warning about 'model welfare'

#568
The concept is based on at least 2 false premises.

I am not aware of a social contract that says we must grant conscious beings rights.

Not aware of a shared, concrete definition of consciousness either, which means no way of deciding whether AI is conscious.

Close to half of us don’t even feel compelled to grant rights to humans for just being humans.

We grant rights to animals, because we love/like them. We enjoy experiencing them. We find them pretty etc. There is a ton of undisputable warm fuzzy.

Humans have rights because they won’t stop being a pain in the back about it. Those that stop lose their rights.

Plain as day for me. Not sure what I am missing or whether I am just a simpleton.

Re: A warning about 'model welfare'

#569

The concept is based on at least 2 false premises. I am not aware of a social contract that says we must grant conscious beings rights. Not aware of a shared, concrete definition of consciousness either, which means no way of deciding whether AI is conscious. Close to half of us don’t even feel compelled to grant rights to humans for just being humans. We grant rights to animals, because we love/like them. We enjoy e…

We grant rights to some animals. Most domesticated animals have very few rights and suffer fates you wouldn't wish upon your worst enemy.

Having rights should not be based upon being conscious/not conscious, but on the ability to suffer. AIs cannot suffer, as far as I'm aware.

Re: A warning about 'model welfare'

#570
post #490

Earlier quoted context omitted.

> But we already have examples that break that rule (us). No, we don't. Where are you reading your research papers?

Please don't insult me.

It's not an insult.

"People are deterministic" is news to me, so I'd really rather like to know which papers claimed that.

Post reply on HN