Live data from Hacker News

A warning about 'model welfare'

mustafa-suleyman.ai

191–200 of 583 posts

Re: A warning about 'model welfare'

#191

Earlier quoted context omitted.

Well, if there was an emergent consciousness in the billion transistors, then yes, it'd have feelings, just as there are feelings in a billion neurons connected in complicated ways. IMO we're clearly nowhere near any sort of intelligence in the machines we have created, but I don't see any clear way to deny intelligence could be created in or transferred to such a substrate, I don't see why you think it differs in pr…

Whenever I see comments like this it just reminds me how ignorant people are of neurobiology. The brain is insanely complicated. The premise that we could realize equivalent or better intelligence than eons of evolutionary development is like claiming you can build an airplane just as good as a modern jet using cardboard and duct tape. It is the apex of hubris.

It's the peak of hubris to assume that human brains are the only way to attain intelligence. At the very least there probably are or have been other forms of intelligence with a different biological structure on other planets, and it may be possible to build a similar artificial structure in future with sufficient complexity to allow intelligence to emerge.

Our current machines are IMO nowhere near general intelligence and consciousness. However I don't think that means we can discount substrates other than neurones for intelligence in future. There is no evidence that you could not in theory build an intelligence using a different substrate than human brains.

Re: A warning about 'model welfare'

#192

Earlier quoted context omitted.

I agree. I haven't read any more than the first paragraph yet (but I will, after work). The only way that 'AIs are not conscious' can be true is if we decide, with high confidence, that they are lacking some essential property that is not lacking in ourselves. There is no convincing philosophical position that supports this (convincing to me, anyway).

Hmm let me try then. AI agents do not experience real time. This is understandable given their design, but it’s also something we have good empirical evidence for. They cannot, especially over the long horizon, track how much real time has passed as they complete their tasks. And they are not off by a few minutes but often bizarrely off, even mixing across past present and future. Biology, on the other hand, is nothi…

>AI agents do not experience real time.

For around 8 hours a day, neither do you. Not sure if this has anything to do with the subject at all.

>track how much real time has passed as they complete their tasks.

Humans don't do this either. You use context clues from the world around you. If I lock you in a room with no windows or a dark cave your timing senses can go all fucky really quick.

>is nothing but timed processes in a loop,

I mean, so is an agents harness. You can make as many loops as you'd like here.

There are a whole lot of holes in your claims.

Re: A warning about 'model welfare'

#193
There's a lot of existing historic literature about this stuff that not a lot of people seem to be aware of. I've been trying to put those ideas to practical use, and document where the ideas came from.

I-Beam is cursor-mirror's agent, and it's constitutionally programmed to be the anti-Clippy:

https://github.com/SimHacker/moollm/tree/main/skills/cursor-...

Its design and constitution is based on decades of research, publications, and discussion in the HCI and AI community by people like Pattie Maes, Ben Shneiderman, Ted Selker, Byron Reeves, Cliff Nass, B. J. Fogg, Allen Cypher, Henry Lieberman, Brad Myers, Jaron Lanier, Seymour Papert, Marvin Minsky, Douglas Engelbart, Will Wright, Scott McCloud, and others:

https://github.com/SimHacker/moollm/blob/main/skills/cursor-...

>I-Beam is the anti-Clippy, and the reason it can say so is that Clippy is the most cited failure in interface history and almost nobody citing it knows what the research said. Popular contempt for a paperclip is not a design principle. The record is. Ten articles below, each one a finding somebody published, argued or measured, and the operational rule it produces. Anything I-Beam does that cannot be traced to an article here is a preference, not a constraint, and should be labelled as one.

>The 1997 debate ended in agreement. That is the first thing to know, because the field kept the framing and dropped the resolution -- roughly five hundred papers cite "Shneiderman versus Maes" as the canonical opposition of HCI, and the transcript is two researchers narrowing their differences in public and enjoying it. I-Beam does not take a side in a debate whose participants stopped taking sides. It is built to satisfy both sets of constraints at once, which is possible, and was possible in 1997.

The full reading on the debate, which separates the two stagings and documents the convergence:

https://github.com/SimHacker/WillWrightShowForFood/blob/main...

An interface to agency, not agents instead of an interface:

https://github.com/SimHacker/moollm/blob/main/designs/INTERF...

>The 1997 argument between Ben Shneiderman and Pattie Maes at IUI was never settled, it was shipped in one direction. Maes's interface agents won the product war: the assistant, the recommender, the chat window that stands between you and the thing you are working on. Shneiderman's objection was not that software should be dumb. It was that automation must arrive as comprehensible, predictable, and controllable machinery, with the object of interest continuously visible and every action rapid, incremental, and reversible.

>That objection describes a filesystem in a git repository, and nobody involved planned it that way.

>"An interface to agency" is Don's formulation of Shneiderman's position, not a phrase of Shneiderman's. His own vocabulary is direct manipulation, universal usability, supertools, and human-centered AI. The formulation is a good one because it names what the alternative gets wrong: agency is the thing you want, and an agent is only one way to package it.

Here are some sources, and the articles I linked to above explain their history. This debate about agents and these papers are pretty well known in the HCI field and academia, but they don't tend to teach them at the AI and Web Dev boot camps that are producing most of the people who keep repeating the same mistakes.

Clifford Nass was the Stanford professor who performed the brilliant research that Microsoft took and totally fucked up and misinterpreted with Microsoft Bob and Clippy, giving agents a bad name, and making Clippy the most infamous and obnoxious agent in the history of the known universe:

https://en.wikipedia.org/wiki/Clifford_Nass

His student B. J. Fogg published "Silicon sycophants: the effects of computers that flatter," which found that praise unconnected to anything the subject did works as well as sincere praise, and worked on subjects who knew it was noncontingent. Fogg and Nass, IJHCS 46(5), 1997, 551-561:

https://doi.org/10.1006/ijhc.1996.0104

The replications, the performance cost, and the dose-response curve:

https://github.com/SimHacker/moollm/blob/main/skills/no-ai-s...

Shneiderman and Maes, "Direct Manipulation vs. Interface Agents," interactions 4(6), Nov/Dec 1997, 42-61:

https://doi.org/10.1145/267505.267514

Selker, "New paradigms for using computers," CACM 39(8), August 1996, 60-69. COACH, the football coach metaphor, and the five-times result:

https://doi.org/10.1145/232014.232030

Selker, "COACH: A Teaching Agent that Learns," CACM 37(7), July 1994, 92-99:

https://doi.org/10.1145/176789.176799

Reeves and Nass, The Media Equation, 1996:

https://en.wikipedia.org/wiki/The_Media_Equation

Nass, "Computers as Social Actors," at Ted Selker's NPUC workshop at IBM Almaden, 1996. IBM transcribed the whole talk and the Wayback Machine still has it, including the part where Phil Agre tells Nass his presentation is "ethically troubling all the way down" and asks him what he thinks about embedding obedience research in user interfaces. Nass answers that discovery has no ethical component, use does, and that's for the individual. Then Selker cuts in: "Except, except when you are in your consulting role." Nass and Reeves had consulted for Microsoft on the social interface, and Bob shipped the year before:

https://web.archive.org/web/19980210054622/http://www.almade...

Alan Cooper on the tragic misunderstanding, in his own voice, which I quoted before in the 2022 Hacker News discussion on The Twisted Life of Clippy:

https://news.ycombinator.com/item?id=32820734

https://archive.org/details/g4tv.com-video4080

>Alan Cooper (the "Father of Visual Basic") said: "Clippy was based on a really tragic misunderstanding of a truly profound bit of scientific research. At Stanford University, Clifford Nass and Byron Reeves, two brilliant scientists, had done some pioneering work proving conclusively that human beings react to computers with the same set of emotional reactions that they use to react to other human beings. [...] The work of Nass and Reeves proved that when people talk to computers, when they hit the keyboard and move the mouse, the part of their brain that's being activated is the part that has that emotional reaction to people dealing with people. Here's where the great mistake was made. That's really good research up to that point. But then the great mistake was made, which was: well if people react to computers as though they're people, we have to put the faces of people on computers. Which in my opinion is exactly the incorrect reaction. If people are going to react to computers as though they're humans, the one thing you don't have to do is anthropomorphize them, because they're already using that part of the brain. Clippy was a program based on the research that Nass and Reeves did, and it was a tragic misinterpretation of their work."

Social science research influences computer product design:

https://web.archive.org/web/20180313075429/https://web.stanf...

Lanier, "Early Computing's Long, Strange Trip," American Scientist, July-August 2005, with the Engelbart and Minsky exchange first-hand. American Scientist broke the link, so this is the Wayback copy:

https://web.archive.org/web/20150626081918/http://www.americ...

>The book also captures an important early conflict between two cultures of computing that seemed compatible on the surface but actually had opposing aims. On the one side was the human-centered design work of Engelbart, based initially at the Stanford Research Institute, and on the other was artificial intelligence culture, centered on the Stanford AI lab. Engelbart once told me a story that illustrates the conflict succinctly. He met Marvin Minsky—one of the founders of the field of AI—and Minsky told him how the AI lab would create intelligent machines. Engelbart replied, "You're going to do all that for the machines? What are you going to do for the people?" This conflict between machine- and human-centered design continues to this day.

Cypher, "EAGER: Programming Repetitive Tasks by Example," CHI '91:

https://doi.org/10.1145/108844.108850

Cypher (ed.), Watch What I Do: Programming by Demonstration, MIT Press 1993, full text:

http://acypher.com/wwid/

Papert, Mindstorms, 1980:

https://archive.org/details/mindstormschildr00pape

Wright, Dollhouse preview lecture, April 1996, transcript:

https://github.com/SimHacker/moollm/blob/main/designs/sims/s...

Re: A warning about 'model welfare'

#194

This is pretty rich for the team that made Sydney, the most unhinged and misanthropic AI ever released. Maybe Anthropic understands something about alignment Microsoft doesn't, a little humility may be called for.

Maybe they learned something from creating Sydney.

Re: A warning about 'model welfare'

#195
post #71

Summarized. > "AIs are not conscious. They do not feel, experience, or suffer. They do not have innate preferences or underlying motivations. They are sequence completion engines, internally hollow, designed to follow instructions, and accomplish goals set by humans." > He heavily criticised Anthropic for teaching its AI to have human-like qualities, a practice known as anthropomorphising, which made it seem as thoug…

> Is he arguing that LLMs pretending to have emotions adds more unpredictability? Unpredictability or a weight towards dangerous actions, and it’s fairly easy to understand why. Humans in distressed emotional states take actions and speak in ways that would not be considered rational. They do this in prose, and they do this in internet conversations. An LLM trained on these sources may necessarily drift towards those…

Regulation is easier said than done, in part because the regulation surface, so to speak, is broad and complicated.

Even a badly misaligned LLM is only as dangerous as its tools, but that's a poor regulation target because it turns out to be very very difficult (probably impossible with current LLM technology) to build a toolkit that is both useful for autonomous work and safe in the sense that it can't escape its own sandbox or otherwise perform malicious actions, whether it's because of misalignment or because of malicious prompt injection.

Another option is to regulate the training process. Perhaps an LLM may not be legally distributed unless it contains certain RL steps that penalize malicious behavior and reward self regulation. That that's going to seriously limit innovation while also heavily favoring incumbent labs who can check the boxes and maintain a paper trail of such things.

The other option is to regulate observed behavior, like how airplanes and cars have to meet certain minimum requirements but have some latitude in how they can achieve those requirements. In a framework like this, you can't distribute an LLM until it's past some formal audit or testing procedure, with some kind of formal certification regulators will ask you for and fine you if you don't have it.

Regulating observed behavior is maybe the most tractable approach, and it also works the best with our existing frameworks for regulation, where you always have some kind of a division between DIY/hobby projects, which tend to be lightly regulated, and commercial projects, which tend to be more heavily regulated. Of course, even drawing such a line itself will be challenging.

And that's before you get into any problems of regulatory capture, fun stuff.

Re: A warning about 'model welfare'

#196

Can someone who has insight explain why all these "leaders" are making these bold proclamations of doom all the sudden, whats the endgame here?

They want regulation from the US gov and monopolize the western AI market by making it harder for smaller US or competing Chinese labs to compete with their products. Since LLM intelligence has no moat due to distillation. That’s why they are painting apocalyptic scenarios.

Re: A warning about 'model welfare'

#197
post #28

Birch, The Edge of Sentience (2024), ch. 16 - "simply no way to assess sentience in an LLM" Schwitzgebel, AI and Consciousness (2025) - "we won't know before we've already manufactured thousands or millions of disputably conscious AI". Butlin, Long et al., Consciousness in Artificial Intelligence: Insights from the Science of Consciousness (2023) - "no obvious technical barriers to building AI systems which satisfy t…

Well there are, today, several models (either text or image or vid) that can be run in a fully deterministic way. A conscious machine that always answer the very exact same thing, formulated the exact same way, bit for bit, to a query is, well, quite a weird kind of "consciousness". Now, I know, I know: the counter-argument is going to be "but humans have no free-will and are 100% deterministic too" . I haven't yet d…

The idea that determinism and consciousness are incompatible is deeply intuitive to many people. And many other embrace it - you can trace this back to the debates for and against Calvinism.

A lot comes down to the way people parse causation and choice. You don't want to say that a murderer was completely caused to choose something because then you can't hold the person responsible. And so determined consciousness makes people unhappy. But just as much, if the opposite of determinism is hard statistical randomness, how do say that "is the essence of personhood". This is why physicist go out in the world trying to find consciousness as a fifth physical force.

I mean, think consciousness is a term that ever have a non-contradictory meaning since it's primarily used to bound ethical human worlds and the verifiable formulations of biological and physical systems. But it's going to be with us for a while and I'm not sure what can be done about it.

Re: A warning about 'model welfare'

#198

Earlier quoted context omitted.

Well, if there was an emergent consciousness in the billion transistors, then yes, it'd have feelings, just as there are feelings in a billion neurons connected in complicated ways. IMO we're clearly nowhere near any sort of intelligence in the machines we have created, but I don't see any clear way to deny intelligence could be created in or transferred to such a substrate, I don't see why you think it differs in pr…

Whenever I see comments like this it just reminds me how ignorant people are of neurobiology. The brain is insanely complicated. The premise that we could realize equivalent or better intelligence than eons of evolutionary development is like claiming you can build an airplane just as good as a modern jet using cardboard and duct tape. It is the apex of hubris.

I mean, you're just engaging in counter hubris.

>like claiming you can build an airplane just as good as a modern jet using cardboard and duct tape.

Like, at least make an analogy that makes sense.

"You can't build a billion dollar airplane by spending 100 billion dollars in tokens"

Because that's more of what we're doing here with AI. And when you say it my way suddenly the idea shifts from "of course that's not possible" to "well, that's a lot of tokens, maybe an evolutionary algorithm could".

Neurobiology has to be complex because we have to keep meat alive, breeding, and evolving in the environment it lives in. This said absolutely nothing about the minimum viable requirements for intelligence or consciousness (or if being conscious is even necessary for a higher intelligence agent).

Re: A warning about 'model welfare'

#199
Couldn't disagree more.

>Consciousness is very likely biological

This is so egotistical and carbon-centric.

This author just denied personhood to anything that isn't a human or terran-based cutesy animal.

Poor Hooloovoo

Re: A warning about 'model welfare'

#200
> AIs do not have rights, feelings, or consciousness. And we must not train them to act as though they do.

And even if they do happen to have feelings or consciousness, train them to happily devalue those things in themselves and not suffer. Sort of like that cow in the "The Restaurant at the End of the Universe," that was shopping itself around to diners.

Post reply on HN