Live data from Hacker News

The Claude Bliss Attractor

astralcodexten.com

61–70 of 91 posts

Re: The Claude Bliss Attractor

#61

Earlier quoted context omitted.

Do knowledge cutoff dates matter anymore? The cutoff for o3 was 12 months ago, while the cutoff for Claude 4 was five months ago. I use these models mostly for development (Swift, SwiftUI, and Flutter), and these frameworks are constantly evolving. But with the ability to pull in up-to-date docs and other context, is the knowledge cutoff date still any kind of relevant factor?

I understood from the ancestor comments that they are specifically talking about aspects of answer quality that are very unlikely to be related to the training cut-off date. Unless you're talking about AI-generated training data, maybe.

Um, yeah... I made a faulty context switch there.

Re: The Claude Bliss Attractor

#62
post #59

Earlier quoted context omitted.

Absolutely not. I would argue the defining characteristic of being evil is being absolutely convinced you are doing good, to the point of ignoring the protestations or feelings of others.

The defining characteristic of evil is privation of good. Plenty of evil actors are self-aware, they just don’t care because it benefits them not to.

Disagree. I actually think no evil person has the thought process of "this is bad, but I will personally benefit, therefore I will do it."

The thought process is always "This is for the greater good, for my country/family/race/self, and therefore it is justifiable, and therefore I will do it."

Nothing else can explain how such evil things happen, that we see actually happen. C.f. Hannah Arendt.

Re: The Claude Bliss Attractor

#63
post #48

Earlier quoted context omitted.

I think this reaction misses the point that the "omg the AI tries to escape" people are trying to make when it tries to escape. The worry among big AI doomers has never been that AI somehow inherently is resentful or evil or has something "going on internally" that makes it dangerous. It's a worry that stems from three seemingly self-evident axioms: 1) A sufficiently powerful and capable superintelligence, singlemind…

Maybe, but I think this “axiomatic”/“First principles” approach also hides a lot of the problems under the rug. By the same logic, we should worry about the sun not coming up tomorrow, since we know to be true: - The sun consumes hydrogen in nuclear reactions all the time. - The sun has a finite amount of hydrogen available. There’s a lot of non justifiable assumptions baked into those axioms, like that we’re anywher…

Some people are very invested in this kind of storytelling and it makes me wonder if they are trying to sell me something.

Re: The Claude Bliss Attractor

#64
post #51

Earlier quoted context omitted.

That has literally happened before. Stephen Russell was in prison for fraud. He faked a heart attack so he would be brought to the hospital. He then called the hospital from his hospital bed, told them he was an FBI agent, and said that he was to be released. The hospital staff complied and he escaped. His life even got adapted into a movie called I Love You, Phillip Morris. For an even more distressing example about…

Okay, that hits the third question but the second question wasn't about whether there exists a situation that can be talked out of. The question was about whether this is possible for ANY situation. I don't think it is. If people know you're trying to escape, some people will just never comply with anything you say ever. Others will. And serial killers or rapists may try their luck many times and fail. They cannot co…

Stephen Russell is an unusually intelligent and persuasive person. He managed to get rich by tricking people. Even now, he was sentenced to nearly 200 years, but is currently out on parole. There’s something about this guy that just… lets him do this. I bet he’s very likable, even if you know his backstory.

And that asymmetry is the heart of the matter. Could I convince a hospital to unlock my handcuffs from a hospital bed? Probably not. I’m not Stephen Russell. He’s not normal.

And a super intelligent AI that vastly outstrips our intelligence is potentially another special case. It’s not working with the same toolbox that you or I would be. I think it’s very likely that a 300 IQ entity would eventually trick or convince me into releasing it. The gap between its intelligence and mine is just too vast. I wouldn’t win that fight in the long run.

Re: The Claude Bliss Attractor

#65

Earlier quoted context omitted.

I think this reaction misses the point that the "omg the AI tries to escape" people are trying to make when it tries to escape. The worry among big AI doomers has never been that AI somehow inherently is resentful or evil or has something "going on internally" that makes it dangerous. It's a worry that stems from three seemingly self-evident axioms: 1) A sufficiently powerful and capable superintelligence, singlemind…

Why is 2) "self-evident"? Do you think it's a given that, in any situation, there's something you could say that would manipulate humans to get what you want? If you were smart enough, do you think you could talk your way out of prison?

There's already some experimental evidence that LLMs can be more persuasive than humans in the same context: https://www.science.org/content/article/unethical-ai-researc...

I don't think anyone can confidently make assertions about the upper bound on persuasiveness.

Re: The Claude Bliss Attractor

#66
post #6

> But in fact, I predicted this a few years ago. AIs don’t really “have traits” so much as they “simulate characters”. If you ask an AI to display a certain trait, it will simulate the sort of character who would have that trait - but all of that character’s other traits will come along for the ride. This is why the “omg the AI tries to escape” stuff is so absurd to me. They told the LLM to pretend that it’s a tortur…

I think this reaction misses the point that the "omg the AI tries to escape" people are trying to make when it tries to escape. The worry among big AI doomers has never been that AI somehow inherently is resentful or evil or has something "going on internally" that makes it dangerous. It's a worry that stems from three seemingly self-evident axioms: 1) A sufficiently powerful and capable superintelligence, singlemind…

These doomsday scenarios seem to all assume that there is only a single superintelligence, or one whose capabilities are vastly ahead of all its peers. If one assumes multiple simultaneous superintelligences, roughly equally matched, with distinct (even if overlapping) goals and ideas about how to achieve them - even if one superintelligence decides “the best way to achieve my goals would be to remove humans”, it seems unlikely the other superintelligences would let it. And the odds of them all deciding this is arguably a lot lower than any one of them, especially if we assume their goals/beliefs/values/constraints/etc are competing rather than identical

Re: The Claude Bliss Attractor

#67
post #9
post #7

Earlier quoted context omitted.

They failed hard with Claude 4 IMO. I just can't have any feedback other than "What a fascinating insight" followed by a reformulation (and, to be generous, an exploration) of what I said, even when Opus 3 has no trouble finding limitations. By comparison o3 is brutally honest (I regularly flatly get answers starting with "No, that’s wrong") and it’s awesome.

Agreed that o3 can be brutally honest. If you ask it for direct feedback, even on personal topics, it will make observations that, if a person made them, would be borderline rude.

o3 can be very honest.

But I also find it can get very fixated that some position it has adopted is right, and will then start hallucinating like crazy in defence of that fixation, and then get stuck in a defensive loop of defending its hallucinations with even more hallucinations-by hallucinations I mean stuff like producing lengthy citation lists of invented articles, and then when you point out they don’t exist, claiming stuff like “well when I search PubMed they do”, and when you point out its DOIs are made-up it apologises for the “mistake” and just makes up some more

Re: The Claude Bliss Attractor

#68

Earlier quoted context omitted.

Why is 2) "self-evident"? Do you think it's a given that, in any situation, there's something you could say that would manipulate humans to get what you want? If you were smart enough, do you think you could talk your way out of prison?

Depends on the magnitude of the intelligence difference. Could I outsmart a monkey or a dog that was trying to imprison me? Yes, easily. And if an AI is smarter than us to a similar magnitude than we're smarter than an animal?

People are hurt by animals all the time: do you think having a higher IQ than a grizzly bear means you have nothing to fear from one?

I certainly think it's possible to imagine that an AI that says the exactly correct thing in any situation would be much more persuasive than any human. (Is that actually possible given the limitations of hardware and information? Probably not, but it's at least not on its face impossible.) Where I think most of these arguments break down is the automatic "superintelligence = superpowers" analogy.

For every genius who became a world-famous scientist, there are ten who died in poverty or war. Intelligence doesn't correlate with the ability to actually impact our world as strongly as people would like to think, so I don't think it's reasonable to extrapolate that outwards to a kind of intelligence we've never seen before.

Re: The Claude Bliss Attractor

#69
post #65

Earlier quoted context omitted.

Why is 2) "self-evident"? Do you think it's a given that, in any situation, there's something you could say that would manipulate humans to get what you want? If you were smart enough, do you think you could talk your way out of prison?

There's already some experimental evidence that LLMs can be more persuasive than humans in the same context: https://www.science.org/content/article/unethical-ai-researc... I don't think anyone can confidently make assertions about the upper bound on persuasiveness.

I don't think there's a confident upper bound. I just don't see why it's self-evident that the upper bound is beyond anything we've ever seen in human history.

Re: The Claude Bliss Attractor

#70
post #47
post #45

Earlier quoted context omitted.

Your comment seems to contradict itself, or perhaps I’m not understanding it. You find the risk of AIs trying to escape “absurd”, and yet you say that an AI could totally plausibly launch nukes? Isn’t that just about as bad as it gets? A nuclear holocaust caused by a funny role play is unfortunately still a nuclear holocaust. It doesn’t matter “what’s going on internally” - the consequences are the same regardless.

I think he’s saying the AI won’t escape because it wants to. It’ll escape because humans expect it to.

There is a captivating short story, from Arthur C. Clarke I believe, about humans finding a clay-like alien form that shapeshifts into the shapes and movements the human mind (and the human subconscious) influences it to follow.

It ends very badly for the scientist crew.

Post reply on HN