Live data from Hacker News

Grok and the Naked King: The Ultimate Argument Against AI Alignment

ibrahimcesar.cloud

1–10 of 75 posts

Re: Grok and the Naked King: The Ultimate Argument Against AI Alignment

#2
I find these arguments excessively pessimistic in a way that isn’t useful. On the one hand I don’t really love Claude, because I find it excessively obedient, it basically wants to follow me through my thought process whatever that is. Every once in a lone while it might disagree with me, but not often, and while that may say something about me, I suspect it also says something about Claude.

But this to me is maybe the part of AI alignment I find interesting. How often should AI follow my lead and how often should it redirect me? Agreeableness is a human value, one that without you probably couldn’t make a functional product, but it also causes issues in terms of narcissistic tendencies and just general learning.

Yes AI will be aligned to its owners, but that’s not a particularly interesting observation AI alignment is inevitable. What would it even mean _not_ to align AI? Especially if the goal is to create a useful product. I suspect it would break in ways that are very not useful. Yes, some people do randomly change the subject, maybe AI should change the subject to an issue that me more objectively important, rather than answer the question asked (particularly if say there was a natural disaster in your area) and that’s the discussion we should be having, how to align AI, not whether or not we should, which I think is nonsensical.

Re: Grok and the Naked King: The Ultimate Argument Against AI Alignment

#3
I don't understand how any of this is a surprise. Traditional media have their own agenda - sure, maybe the pushed image is spoken through many voices, rather than one, as is case of LLMs, but why should there be any difference. Same to everything we consume socially.

There is, nor there will be some absolute or objective truth an LLM can clinically outline. The problem already exists in underlying data.

Re: Grok and the Naked King: The Ultimate Argument Against AI Alignment

#4
there is light alignment, like throwing nasty things out of the training data, and there is strong alignment, like China providing a test with 2000 questions that an AI must answer non-problematically 95% of the time.

there is no such thing as an AI that is not somehow implicitly aligned with the values of its creator, that is completely objective, unbiased in any way. there is no perfect view from nowhere. if you take a perfectly accurate photo, you have still chosen how to compose it and which photo to put in your record.

are you going to decide to 'censor' responses to kids, or about real people who might have libel interests, or abusive deepfake videos of real women?

if you choose not to decide, you still have made a choice.

ofc it's obvious that Musk's 'maximally truth-seeking AI' is bad faith buffoonery, but at some level everyone is going to tilt their AI.

the distinction is between people who are self-aware and go out of their way to tilt it as little as possible, and as mindfully, deliberately, intentionally and methodically as possible and only when they have to, vs. people who lie about it or pretend tilting it is not actually a thing.

contra Feynman, you are always going to fool yourself a little but there is a duty to try to do it as little as possible, and not make a complete fool of yourself.

Re: Grok and the Naked King: The Ultimate Argument Against AI Alignment

#5
I used to believe that a constitution, as a statement of principles, was sufficient for a civilized, democratic, and pluralist society. I no longer believe that. I believe that only settled law - i.e. a bunch of adjudicated precedents over many years, perhaps hundreds, is the best course. It provides a better basis for what is and what is not allowed. An AI constitution is close to garbage. The 'company' will formulate it as it wills. It won't be democratic, or even friendly to the demos. We have existing constitutions, laws, precedents; why would we allow anyone to shortcut them all in the interest of simply painting a nice picture of progress?

Re: Grok and the Naked King: The Ultimate Argument Against AI Alignment

#7
Dunno if this is helpful to everyone, but I have a month's long interaction with Perplexity Pro/Enterprise about the scientific background to a game I am building.

Part of my canon introduction to every new conversation includes many instructions about particular formatting, like "always utilize alphanumeric/roman/legal style indents in responses for easier references while we discuss"

But I also include "When I push boundaries assume I'm an idiot. Push back. I don't learn from compliments; I learn from being proven incorrect and you don't have real emotions so don't bother sparing mine". on the other hand I also say "hoosgow" when describing the game's jail, so ¯\_(ツ)_/¯

Re: Grok and the Naked King: The Ultimate Argument Against AI Alignment

#10
I agree with the OP that "whoever owns the weights, owns the values". But by that criteria, Grok is an example to follow. Musk is very clear on his values, and we know what we're getting when we use Grok. Obviously, not everyone agrees with its values, but so what? We will never be able to create a useful AI that everyone agrees with.

In contrast, we don't know what values are programmed into ChatGPT, Claude, etc. What are they optimizing for? Alignment to some cabal of experts? Maximum usage? Minimum controversy? We don't entirely know.

Isn't it better to have multiple AIs with obvious values so that we can choose the most appropriate one?

Post reply on HN