Live data from Hacker News

The Claude Bliss Attractor

astralcodexten.com

81–90 of 91 posts

Re: The Claude Bliss Attractor

#81

Earlier quoted context omitted.

If someone who is so good at manipulation their life is adapted into a movie still ends up serving decades behind bars, isn't that actually a pretty good indication that maxing out Speech doesn't give you superpowers? AI that's as good as a persuasive human at persuasion is clearly impactful, but I certainly don't see it as self-evident that you can just keep drawing the line out until you end up with 200 IQ AI that…

In the context of the topic (could a rogue super intelligence break out), I don’t really see how that’s relevant. Clearly someone who is clever enough has an advantage at breaking out. As for the bit about how limited it is, do you remember the Rowhammer attack? https://en.m.wikipedia.org/wiki/Row_hammer This is exactly the kind of thing I’d worry about a super intelligence being able to discover about the hardware i…

Advantage, sure. I just don't think that advantage is particularly meaningful in situations a human has virtually no chance of escaping. Humans also have a lot of their own advantages. How is a chatbot supposed to cross an air gap unless you assume it has what I consider unrealistic levels of persuasion?

I think you also have to consider that AI with superpowers is not going to materialize overnight. If superintelligent AI is on the horizon, the first such AI will be comparable to very capable humans (who do not have the ability to talk their way into nuclear launch codes or out of decades-long prison sentences at will). Energy costs will still be tremendous, and just keeping the system going will require enormous levels of human cooperation. The world will change a lot in that kind of scenario, and I don't know how reasonable it is to claim anything more than the observation of potential risks in a world so different from the one we know.

Is it possible that search ends up doing as much for persuasion as it does for chess, superintelligent AI happens relatively soon, and it doesn't have prohibitive energy costs such that escape is a realistic scenario? I suppose? Is any of that obvious or even likely? I wouldn't say so.

Re: The Claude Bliss Attractor

#82

Earlier quoted context omitted.

In the context of the topic (could a rogue super intelligence break out), I don’t really see how that’s relevant. Clearly someone who is clever enough has an advantage at breaking out. As for the bit about how limited it is, do you remember the Rowhammer attack? https://en.m.wikipedia.org/wiki/Row_hammer This is exactly the kind of thing I’d worry about a super intelligence being able to discover about the hardware i…

Advantage, sure. I just don't think that advantage is particularly meaningful in situations a human has virtually no chance of escaping. Humans also have a lot of their own advantages. How is a chatbot supposed to cross an air gap unless you assume it has what I consider unrealistic levels of persuasion? I think you also have to consider that AI with superpowers is not going to materialize overnight. If superintellig…

Oh, I don’t think we’re on the verge of super-intelligence. I don’t think LLMs are a pathway to super-intelligence, and I think most of the talk of super-intelligence right now is marketing fluff.

In terms of crossing an air gap, that really depends. For example, are you aware that researchers can pretty reliably figure out what someone typed just by the sound of the keys being pressed?

Or how about that team who developed a code analysis tool to detect errors, and they ran Tim-sort through it, and the tool said that Tim-sort had a bug where certain pathological inputs would cause it to crash? The researchers assumed their tool was incorrect because Tim-sort is so widely used (it’s the default sort algorithm in Python and Android, for example). But they decided to try it out to see what would happen, and sure enough, they could hard-crash the Python interpreter. No one realized this bug had been there the whole time.

Or various image codec bugs over the years that have allowed a device to be compromised just by viewing an image?

There are some weird bugs out there. Are we certain that there’s no way a computer could detect variances in the timings or voltages happening within it to act as a WiFi antenna or something like that? We’ve found some weird shit over the years! And a super-intelligence that’s vastly smarter than us is way more likely to find it than we are.

Basically, no, I don’t trust the air gap with a sufficiently advanced super intelligence. I think there are things we don’t know that we don’t know, and a super-intelligence would spot them way before we would. There are probably a hundred more Rowhammer attacks out there waiting to be discovered. Are we sure none of them could exchange data with a nearby device? I’m not.

Re: The Claude Bliss Attractor

#83

This thread is getting at a key point, and `roxolotl` is right on the money about roleplaying. The anxiety about AI 'wants' or 'bliss' often conflates two different processes: forward and backward propagation. Everything we see in a chat is the forward pass. It's just the network running its weights, playing back a learned function based on the prompt. It's an echo, not a live thought. If any form of qualia or genuin…

Interesting idea! Could you tell me more about how you arrive at the conclusion that forward passes cannot produce qualia?

Re: The Claude Bliss Attractor

#84

Earlier quoted context omitted.

Why is 2) "self-evident"? Do you think it's a given that, in any situation, there's something you could say that would manipulate humans to get what you want? If you were smart enough, do you think you could talk your way out of prison?

That has literally happened before. Stephen Russell was in prison for fraud. He faked a heart attack so he would be brought to the hospital. He then called the hospital from his hospital bed, told them he was an FBI agent, and said that he was to be released. The hospital staff complied and he escaped. His life even got adapted into a movie called I Love You, Phillip Morris. For an even more distressing example about…

> Stephen Russell was in prison for fraud. He faked a heart attack so he would be brought to the hospital

According to Wikipedia he wasn't in prison, he was attempting to con someone at the time and they got suspicious. He pretended to be an FBI agent because he was on security watch. Still impressive, but not as impressive as actually escaping from prison that way.

Re: The Claude Bliss Attractor

#85
post #2

> Anthropic deliberately gave Claude a male name to buck the trend of female AI assistants (Siri, Alexa, etc). In France, the name Claude is given to males and females.

It’s actually really sexist that when the first truly intelligent AI emerge people would now give them male names.

Yes. It is pity that common male name of o5-pro-mini is used to refer to AGI instead of widespread female name Deepseek R7.

Re: The Claude Bliss Attractor

#86

This thread is getting at a key point, and `roxolotl` is right on the money about roleplaying. The anxiety about AI 'wants' or 'bliss' often conflates two different processes: forward and backward propagation. Everything we see in a chat is the forward pass. It's just the network running its weights, playing back a learned function based on the prompt. It's an echo, not a live thought. If any form of qualia or genuin…

Interesting idea! Could you tell me more about how you arrive at the conclusion that forward passes cannot produce qualia?

https://github.com/dmf-archive/IPWT

Re: The Claude Bliss Attractor

#87
post #62

Earlier quoted context omitted.

The defining characteristic of evil is privation of good. Plenty of evil actors are self-aware, they just don’t care because it benefits them not to.

Disagree. I actually think no evil person has the thought process of "this is bad, but I will personally benefit, therefore I will do it." The thought process is always "This is for the greater good, for my country/family/race/self, and therefore it is justifiable, and therefore I will do it." Nothing else can explain how such evil things happen, that we see actually happen. C.f. Hannah Arendt.

You read too many comic books.

You think your average, low-level gang member, willing to murder a random person for personal status, thinks what they're doing is "for the greater good"? They do what they do because they value human life less than they value their own material benefit.

Most evil is banal and pedestrian. It is selfish, short-sighted, and destructive.

"Privation of good" explains all evil.

Re: The Claude Bliss Attractor

#88

Earlier quoted context omitted.

I think this reaction misses the point that the "omg the AI tries to escape" people are trying to make when it tries to escape. The worry among big AI doomers has never been that AI somehow inherently is resentful or evil or has something "going on internally" that makes it dangerous. It's a worry that stems from three seemingly self-evident axioms: 1) A sufficiently powerful and capable superintelligence, singlemind…

These doomsday scenarios seem to all assume that there is only a single superintelligence, or one whose capabilities are vastly ahead of all its peers. If one assumes multiple simultaneous superintelligences, roughly equally matched, with distinct (even if overlapping) goals and ideas about how to achieve them - even if one superintelligence decides “the best way to achieve my goals would be to remove humans”, it see…

> These doomsday scenarios seem to all assume that there is only a single superintelligence, or one whose capabilities are vastly ahead of all its peers

A lot of scenarios posit that this is basically guaranteed to happen the very first time we have an ASI smart enough to bootstrap itself into greater intelligence, unless we're basically perfect in how we align its goals, so yes.

LLMs seem like the best case world for avoiding this scenario (improvement requires exponentially scaling resources compared to inference). That said, it is by no means guaranteed that LLMs will remain the SotA for AI.

> “the best way to achieve my goals would be to remove humans”, it seems unlikely the other superintelligences would let it.

Why not? Seriously, why wouldn't the other superintelligences let it? There's no reason to assume that, by default, ASI would be invested in the survival of the human race in a way we would prefer. More than likely they will just be as laser focussed on their own goals.

The whole point is that it's very difficult to design an AI that has "defending what humans think is right" as a terminal value. They basically all try and find loopholes. The only way to make safety its number one priority is to dial it up so high that everyone complains it's a puritan - and then you get scenarios where it's telling schizophrenics to go off their meds because it's so agreeable.

Unless you're saying that they would fight off the other ASIs gaining power because they themselves want it, but I fail to see how being stuck in a turf war between two ASIs is at all an improvement.

Re: The Claude Bliss Attractor

#89
post #48

Earlier quoted context omitted.

I think this reaction misses the point that the "omg the AI tries to escape" people are trying to make when it tries to escape. The worry among big AI doomers has never been that AI somehow inherently is resentful or evil or has something "going on internally" that makes it dangerous. It's a worry that stems from three seemingly self-evident axioms: 1) A sufficiently powerful and capable superintelligence, singlemind…

Maybe, but I think this “axiomatic”/“First principles” approach also hides a lot of the problems under the rug. By the same logic, we should worry about the sun not coming up tomorrow, since we know to be true: - The sun consumes hydrogen in nuclear reactions all the time. - The sun has a finite amount of hydrogen available. There’s a lot of non justifiable assumptions baked into those axioms, like that we’re anywher…

> There’s a lot of non justifiable assumptions baked into those axioms, like that we’re anywhere close to superintelligence or the sun running out of hydrogen.

Nowhere in my post did I imply a timeline for this. The first predicate of the argument is when we eventually do develop ASI. You can make plans for that, same as you can make plans for when the sun eventually runs out of hydrogen. The difference is, we can look at the sun and say that at the rate it's burning, we've got maybe billions of years, and then look at AI improvements, extrapolate, and assume we've got less than 100 years. If we had any indication the sun was going to nova in 100 years, people would be way more worried.

Re: The Claude Bliss Attractor

#90

Earlier quoted context omitted.

I think this reaction misses the point that the "omg the AI tries to escape" people are trying to make when it tries to escape. The worry among big AI doomers has never been that AI somehow inherently is resentful or evil or has something "going on internally" that makes it dangerous. It's a worry that stems from three seemingly self-evident axioms: 1) A sufficiently powerful and capable superintelligence, singlemind…

I think the "escape the box" explanation misses the point, if anything. The same problem has been super visible in RL for a long time, and it's basically like complaining that water tries to "escape a box" (seek a lowest point). Give it rules and it will often violate their spirit ("cheat"). This doesn't imply malice or sentience.

I think the second sentence I wrote makes it clear that whether it's malicious is irrelevant. Water doesn't have to be malicious for you to observe that we shouldn't be building dams stretching into space next to populated cities when our technology to make dams higher is moving faster than our tech to make them stronger.

And it can still make you nervous when the tiny dams in unpopulated areas start breaking when they're made too tall, even though everyone is telling you that "of course they broke, they're much thinner than city dams" (while they keep building the city dams taller but not thicker)

Post reply on HN