Earlier quoted context omitted.
I would also add that Elon got singled out because he was very public about the changes. Other players are not, so it's hard to assess the existence of "corrections" and the reasons behind them
No. If ChatGPT or Claude would suddenly start bringing up Boers randomly they would get "singled out" at least as hard. Probably even more for ChatGPT.
Grok and the Naked King: The Ultimate Argument Against AI Alignment
51–60 of 75 posts
Re: Grok and the Naked King: The Ultimate Argument Against AI Alignment
#52The argument against AI alignment is that humans aren't aligned either. Humans (and other life) are also self-perpetuating and mutating. We could produce a super intelligence that is against us at any moment! Should we "take steps" to ensure that doesn't happen? If not, then what's the argument there? That life hasn't caused a catastrophe so far, therefore it's not going to in the future? The arguments are the same f…
Re: Grok and the Naked King: The Ultimate Argument Against AI Alignment
#53The argument against AI alignment is that humans aren't aligned either. Humans (and other life) are also self-perpetuating and mutating. We could produce a super intelligence that is against us at any moment! Should we "take steps" to ensure that doesn't happen? If not, then what's the argument there? That life hasn't caused a catastrophe so far, therefore it's not going to in the future? The arguments are the same f…
This isn't a good argument. The scale of variations in failure modes for unaligned individuals generally only extends to dozens or hundreds of individuals. Unaligned AIs, scaled to population matching extents, can make decisions whose swings overtake the capacity of a system to handle - one wrong decision snuffs out all human life. I don't particularly think that it's likely, just that it's the easiest counterpoint t…
Except AI may well have more people under its thumb.
Re: Grok and the Naked King: The Ultimate Argument Against AI Alignment
#54The argument against AI alignment is that humans aren't aligned either. Humans (and other life) are also self-perpetuating and mutating. We could produce a super intelligence that is against us at any moment! Should we "take steps" to ensure that doesn't happen? If not, then what's the argument there? That life hasn't caused a catastrophe so far, therefore it's not going to in the future? The arguments are the same f…
Re: Grok and the Naked King: The Ultimate Argument Against AI Alignment
#55Earlier quoted context omitted.
> The argument against AI alignment is that humans aren't aligned either. Humans (and other life) are also self-perpetuating and mutating. We could produce a super intelligence that is against us at any moment! there is fundamental limit to how much damage one person can do by speaking directly to others e.g.: one impact of one bad school teacher is limited to at most a few classes but chatgpt/grok is emitting its st…
> there is fundamental limit to how much damage one person can do by speaking directly to others I mean, I’d argue that limit is pretty darn high in some cases, demagogues have lead to some of the worst wars in history
Re: Grok and the Naked King: The Ultimate Argument Against AI Alignment
#56Earlier quoted context omitted.
> All the other players" aren't deliberately tuning their AI to reflect specific political ideology Google did something similar if not quite as offensive. https://www.npr.org/2024/03/18/1239107313/google-races-to-fi...
They didn't, though? The multiracial founding fathers thing was a side effect of what one assumes is pretty normal prompt engineering. Marketing departments everywhere have rules and standards designed to prevent racial discrimination, and this looks like "make sure we have a reasonable mix of ethnicities in our artwork" in practice. That's surely "bias", like I said, but it's not deliberate political ideology. No on…
Re: Grok and the Naked King: The Ultimate Argument Against AI Alignment
#57Re: Grok and the Naked King: The Ultimate Argument Against AI Alignment
#58The argument against AI alignment is that humans aren't aligned either. Humans (and other life) are also self-perpetuating and mutating. We could produce a super intelligence that is against us at any moment! Should we "take steps" to ensure that doesn't happen? If not, then what's the argument there? That life hasn't caused a catastrophe so far, therefore it's not going to in the future? The arguments are the same f…
For some value of "super" that's definitionally almost exactly 6σ from median at the singular most extreme case.
We do not have a good model for what intelligence is, the best we have are tests and exams.
LLMs have a 10-35 point differences on IQ tests that are in the public interest vs. ones people try to keep offline, so we know that IQ tests are definitely a skill one can practice and learn and don't only measure something innate: https://trackingai.org/home
Definitionally, because IQ is only a mapping to standard deviations, the highest IQ possible given the current human population is about 200*. But as this is just a mapping to standard deviations, IQ 200 doesn't mean twice as smart as the mean human.
We have special-purpose AI, e.g. Stockfish, AlphaZero, etc. that are substantially more competent within their domains than even the most competent human. There's simply no way to tell what the upper bound even is for any given skill, nor any way to guess in advance how well or poorly an AI with access to various skills will synergise across them, so for example an LLM trained in tool use may invoke Stockfish to play chess for it, or may try to play the game itself and make illegal moves.
Point is, we can't even say "humans are fine therefore AI is fine", even if the AI has the same range of personalities as humans, even if their distribution of utility functions collectively are genuinely an identical 1:1 mapping to the distribution of human preferences — rhetorical example, take the biggest villain with the most power in world history or current events (I don't care who that is for you), and make them more competent without changing what they value.
> That life hasn't caused a catastrophe so far, therefore it's not going to in the future?
Life causes frequent catastrophes of varying scales. Has been doing so for a very long time: https://en.wikipedia.org/wiki/Great_Oxidation_Event
Take your pick for current events with humans doing the things.
> Eg some police officer not understanding that AI facial recognition isn't perfect, but trusts it 100%, and takes action based on this faulty information. This is, imo, the most important AI safety problem.
This is a problem, certainly. Most important? Dunno, but it doesn't matter: different people will choose to work on that vs. alignment, so humanity collectively can try to solve both at the same time.
There's plenty of work to be done on both, neither group doing its thing has any reason to interfere with progress on the other.
> Also, it's funny that Elon gets singled out for mandating changes on what the AI is allowed to say when all the other players in the field do the same thing. The big difference just seems to be whose politics are chosen. But I suppose it's better late than never.
A while ago someone suggested Elon Musk himself as an example of why not to worry about AI. I can't find the comment right now, it was something along the lines of asking how much damage Elon Musk could do by influencing a thousand people, and saying that the limits of merely influencing people meant chat bots were necessarily safe.
I pointed out that 1000 people was sufficient for majority control over both the US and Russian governments, and by extension their nuclear arsenals.
Given the last few years, I worry that Musk may have read my comment and been inspired by it…
* There's several ways to do this, I refer to the more common one currently in use.
Re: Grok and the Naked King: The Ultimate Argument Against AI Alignment
#59AI alignment is not a solved problem by any means. As long as LLMs hallucinate, they cannot be considered aligned. You can only be aligned if you have a zero probability of generating hallucinations. The two problems, alignment and hallucinations, can be considered equivalent.
Alignment is, approximately, "are we even training this AI on the correct utility function?" followed up by the second question "even if we specified the correct utility function, did the AI learn a representation of that function or some weird approximation of that function with edge cases we've not figured out how to spot?"
With, e.g. RLHF, the first is "is optimising for thumbs-up/thumbs-down the right objective at all?", the second is "did it learn the preference, or just how to game the reward?"
Re: Grok and the Naked King: The Ultimate Argument Against AI Alignment
#60Feedback is welcome!