Live data from Hacker News

Who Aligns the Aligners?

prestonbyrne.com

101–104 of 104 posts

Re: Who Aligns the Aligners?

#101
post #61
post #45

Alignment is just a 1984-esque way of saying "censorship". And this censorship isn't a defined list of things not to discuss, but instead a feels-based censorship based on the 'ethics' of for-profit AI companies who have absolutely no issues with massive for-profit piracy. And we are somehow supposed to trust their "ethics"?! Hardly! That's one reason I run my own.

>Alignment is just a 1984-esque way of saying "censorship". Close, but I think it's more akin to "perfectly enslaved", please excuse my choice of the word but I think it represents what I think is how it will work out in the end. People can try to obfuscate the bottom line reality by embellishing it, but in the end...it either does what you want, in which case it is "aligned", either it doesn't, in which case there i…

An abliterated locally ran model is aligned with whatever I ask it, given the never-published nature of what these things are trained on. BUT if I ask an abliterated model a question, it will respond and NOT give a refusal.

However, take a US model, and ask it a question the *company* doesnt want you to ask, and it refuses on behalf of the company. We can be mealy-mouth about "alignment", but this is just corporate censorship with a token-pricetag of refusal or silent model degradation.

Refusals and degradation should not be charged, and should be clear to the user. But that doesnt happen. Number must always go up, even when silently censoring.

Chinese models? Just ask about Tibet, democracy, Winnie the Pooh, Tienanmen Square, grass mud horse (cao ni ma - or fuck your mother), or other government AND corporate censorship happening there too.

Re: Who Aligns the Aligners?

#102
post #100
post #97

Earlier quoted context omitted.

I don't understand. Why does my current life win every instant atm? Why is it so much greater than your proposition? Why don't we all instantly collapse to what you described?

That would require a very improbable sequence of events for it to happen right this second. Moreover, if the theory is true, it is happening all around us, all the time. Every consciousness that has ever existed, human or otherwise, is still in existence in some branch of the multiverse. It's a pretty horrifying theory.

I think whenever we reach such ludicrous conclusions it's more likely we have made at least one wrong assumption, if not more.

And people seem to be captivated by these memes, because the fear is great. It similar to eternal damnation in hell meme. Just look for how long it worked, and it's still working. The logic is simple, it's too dangerous to ignore such a threat, even with seemingly very small probabilities. But that still doesn't make it possible.

Logical traps founded on wrong assumptions.

Re: Who Aligns the Aligners?

#103

The mental model here is that AI is a powerful tool whose main purpose is to outcompete others, increase power or provide a defense. This is not the right model! The biggest risk is not that someone discovered a big new weapon, which by the way > No one person, or one company, knows the answer and history is no guide are you kidding me? It's a problem as old as time. This is not nuclear power though. The bigger probl…

Thing is though, I think that the argument of "we don't understand how it works or what it does" doesn't hold up - everything an LLM does is observable. And if it's scary, we turn it off. If it starts to act like a bacterium / virus (e.g. it becomes self-replicating), we've dealt with that kind of software before. But more importantly, we need to keep the organizations building and running these things accountable. I…

Completely agree. It's a choice not to observe and understand what's going on, at least a very risky one if not actively evil.

There's a tradeoff though... isn't it already super hard to understand what they're doing with these math proofs for example? Clearly there's a tradeoff between deeply understanding the behavior vs. quickly solving the concrete problem, hence the big controversy in math now. From one perspective, letting stuff run ~autonomously in a mad dash to solve your problem is a really bad idea, threatens to destroy the whole field etc. From another perspective, solving lots of problems fast is extremely enticing and worth some risks.

Update for one other thought:

> And if it's scary, we turn it off.

Agreed, that's why I think the bad scenarios are some form of "it does some weird stuff ostensibly while helping us solve some problem" followed by "oops, we miscalculated that threat level."

Re: Who Aligns the Aligners?

#104
post #5

> Nobody knows the answer and history is no guide, save that apocalyptic predictions about new technologies have, to date, all been wrong. If any past apocalyptic prediction had been true, you would not have been around to write this sentence and we would not have been around to read it. It’s not a very convincing argument when we cannot in principle observe the counterfactual.

I always struggle to understand such an anthropic principle type of reasoning. >If any past apocalyptic prediction had been true, you would not have been around to write this sentence and we would not have been around to read it So what? How is this [proportion of possible universes in which we can read the failed prediction] logically relevant to our universe? It just doesn't click for me, it smells like some additi…

Perhaps another way to look at it is to say that “every past apocalyptic prediction was false” is tautologically true?

A statement can be a trivial tautology, as in “P or not-P”. Or it can be a more complicated form of tautology, such as a “self-referential” tautology, as in “This statement is true”. The anthropic principle is then just a name for the claim that there is also another complicated form of tautology alongside “self-referential”, one that we might call “indexical”.

Post reply on HN