Live data from Hacker News

Three Inverse Laws of AI

susam.net

261–270 of 388 posts

Re: Three Inverse Laws of AI

#261
post #192

Earlier quoted context omitted.

Because there are countless instances in the training material where a bank robber scopes out the security cameras.

What's an example then, you can think of, of a question where a human could infer intent but an LLM couldn't?

Just today I asked Claude Code to generate migrations for a change, and instead of running the createMigration script it generated the file itself, including the header that says

  // This file was generated with 'npm run createMigrations' do not edit it
When I asked why it tried doing that instead of calling the createMigrations script, it told me it was faster to do it this way. When I asked you why it wrote the header saying it was auto-generated with a script, it told me because all the other files in the migrations folder start with that header.

Opus 4.7 xhigh by the way

Re: Three Inverse Laws of AI

#262

Earlier quoted context omitted.

We have invented a new tool that can cause great harm. Do you see any value whatsoever in promulgating safety guidelines for humans to use the tool without hurting themselves or others? Do you not own any power tools?

I think in order for "AI safety" to be achievable and effective, we need to have a shared agreement on what "safety" means. Recently, the word has been overloaded to mean all sorts of things and used to justify run-of-the-mill censorship (nothing to do with safety). Safety should go back to being narrowly defined in terms of reducing / preventing physical injury. Safety is not "don't use swear words." Safety is not "…

Okay. What's your easy to adopt, easy to understand replacement word for "Safety" in this case?

Re: Three Inverse Laws of AI

#263

> Humans must not anthropomorphise AI systems. Can someone explain why this is a bad thing, while at the same time it's a good thing to say stuff like "put a computer to sleep", "hibernate", "killing" processes, processes having "child" processes, "reaping", "what does the error say ?", "touch", etc? To me that's just language, and humans just using casual language.

The difference is never before has the presentation of a computer and its capabilities made the person on the other end decide "Wow, this is like talking to a real person. I'm gonna date this computer"

Re: Three Inverse Laws of AI

#264
post #6

I strongly disagree with this framing. It's patently insane to demand that humans alter their behavior to accommodate the foibles of mere machines, and it simply won't work in the majority of cases. Humans WILL anthropomorphize the AI, humans WILL blindly trust their outputs, and humans WILL defer responsibility to them. Asimov's laws of robotics are flawed too, of course. There is no finite set of rules that can con…

I believe "AI safety" is a form of pulling up the ladder, or regulatory market capture.

Re: Three Inverse Laws of AI

#265
post #99

Earlier quoted context omitted.

> So this isn't "accommodating foibles" with the machine, it's protecting ourselves from an exploit of a human vulnerability: we subconsciously tend to infer intent, understanding, judgment, emotions, moral agency, etc. to LLMs. Right, I'm saying that this framing is backwards. It's not that poor little humans are vulnerable and we need to protect ourselves on an individual level, we need to make it illegal and socia…

Ah, I see, you are not American. In the US we don't have the luxury of believing our governments will act in the interests of the voters.

I had a similar thought, that parent commenter sounded like they were in Canada or something. Interesting that their solution is to impose constraints on technological process, rather than finding novel ways to elevate individual and collective human functioning in spite of our limitations. Ironically it's his view that is more anti-human

Re: Three Inverse Laws of AI

#266

I been using codex heavily for the past 6 months and I've observed myself going through different types of emotions. Even now, when it does a sloppy job, I still feel emotion, even while it is just a neutral statistical response, its hard to separate natural human instincts. I often wish I could reach through the screen and give him a good shake. Sometimes I want to thank him but then cannot due to scarcity of weekly…

Consider how sailors lovingly refer to their craft as “she”. My vague sense is that society views this as a positive.

I definitely do not feel codex gives off feminine energy

it feels as frustrating as talking to a junior dev from a decade ago

claude felt more feminine

Re: Three Inverse Laws of AI

#267
post #203

Earlier quoted context omitted.

If you want to convince yourself that they can infer intent despite the fundamental limitations of the systems literally not permitting it then you can be my guest. Faking it is fine, sure, until it can’t fake it anymore. Leading the question towards the intended result is very much what I mean: we intrinsically want them to succeed so we prime them to reflect what we want to see. This is literally no different than…

What is fundamental to LLM's that make it impossible for them to infer intent? All the limitations you are describing with respect to LLM's are the same as humans. Would a human tripping up on an ambiguously worded question mean they are always just faking their thinking?

“We see emotion.”—We do not see facial contortions and make inferences from them … to joy, grief, boredom. We describe a face immediately as sad, radiant, bored, even when we are unable to give any other description of the features." (Wittgenstein)

Re: Three Inverse Laws of AI

#268
post #6

I strongly disagree with this framing. It's patently insane to demand that humans alter their behavior to accommodate the foibles of mere machines, and it simply won't work in the majority of cases. Humans WILL anthropomorphize the AI, humans WILL blindly trust their outputs, and humans WILL defer responsibility to them. Asimov's laws of robotics are flawed too, of course. There is no finite set of rules that can con…

>It's patently insane to demand that humans alter their behavior to accommodate the foibles of mere machines

programers have been doing exactly this for long time.

Re: Three Inverse Laws of AI

#269
post #6

I strongly disagree with this framing. It's patently insane to demand that humans alter their behavior to accommodate the foibles of mere machines, and it simply won't work in the majority of cases. Humans WILL anthropomorphize the AI, humans WILL blindly trust their outputs, and humans WILL defer responsibility to them. Asimov's laws of robotics are flawed too, of course. There is no finite set of rules that can con…

  > It's patently insane to demand that humans alter their behavior to accommodate the foibles of mere machines
I don't think it's insane, we do it all the time. Most tools require training to use properly. Including tools that people use every day and think are intuitive. Use the can opener as an example (I'll leave it for you all to google and then argue in the comments).

The difference here is that this tool is thrust upon us. In that sense I agree with you that the burden of proper usage is pushed onto the user rather than incorporated into the design of the tool. A niche specific tool can have whatever complex training and usage it wants.

But a general access and generally available tool doesn't have the luxury of allowing for inane usage. LLMs and Agents are poorly designed, and at every level of the pipeline. They're so poorly designed that it's incredibly difficult to use them properly and I'll generally agree with you that the rules the author presents aren't going to stick. The LLM is designed to encourage anthropomorphization. Usage highly encourages natural language, which in turn will cause anthropomorphism. The RLHF tuning optimizes human preference which does the same thing as well as envisaged behaviors like deception and manipulation along with truthful answering (those results are not in contention even if they seem so at first glance).

But I also understand the author's motivation. Truth is unless you're going full luddite you're going to be interacting with these machines. Truth is the ones designing them don't give a shit about proper usage, they care more about if humans believe the responses are accurate and meaningful more then they care if the responses are accurate and meaningful[0]. So it's fucked up, but we are in a position where we're effectively forced to deal with this.

So really, I agree with you that this is insane.

> I don't have a proof, but I believe that "AI safety" is inherently impossible, a contradiction of terms

To paraphrase my namesake, there's no axiomatic system that is entirely self consistent.

Though safety and security is rarely about ensuring all edge cases are impossible, but rather bounding. E.g. all passwords are hackable, but the failure mode is bound such that it is effectively impossible to crack, but not technically. (And quantum algorithms do show how some of the assumptions break down with a paradigm shift. What was reasonable before no longer is)

[0] this is part of a larger conversation where the economy is set up such that people who make things are not encouraged to make those things better. I specifically am avoiding the word "product" because the "product" is no longer the thing being built, it's the share holder value. Just like how TV's don't care much about making the physical device better but care much more about their spyware and ads. Or well... just look at Microsoft if you need a few hundred examples

Re: Three Inverse Laws of AI

#270

Earlier quoted context omitted.

One, a placebo does not need to be given blindly. A sugar pill is a placebo, even if the recipient knows about it. An actual definition: "A placebo is an inactive substance (like a sugar pill) or procedure (like sham surgery) with no intrinsic therapeutic value, designed to look identical to real treatment." No mention of the user's belief. Two, real hard data proves that the placebo effect remains (albeit reduced) e…

In psychology, the two main hypotheses of the placebo effect are expectancy theory and classical conditioning.[70] In 1985, Irving Kirsch hypothesized that placebo effects are produced by the self-fulfilling effects of response expectancies, in which the belief that one will feel different leads a person to actually feel different.[71] According to this theory, the belief that one has received an active treatment can…

>”Wouldn't want some placebo fiend to O.D.”

We should be more worried about the rise of placebo resistant bacteria.

Post reply on HN