Live data from Hacker News

An Alien Mind

openai.com

391–400 of 490 posts

Re: An Alien Mind

#391

Earlier quoted context omitted.

Would the recent AI-aided math proofs count as breakthroughs?

Prepare for moving goalposts. In 50 years people will still doubt that AI can create anything novel and worthwhile, while they are going to rely mostly on the things that didn't exist before AI, in nutrition, medicine, technology, communication, entertainment. They will see them as, normal, common and simple extensions of the previous developments, pushed mindlessly a bit forward by stochastic parrots.

Claiming that AI will be valuable in 50 years is you moving the goalposts back 50 years. It’s a strawman argument, no one is saying it won’t be valuable in 50 years

Re: An Alien Mind

#392
I find it odd that discussions of alignment don't mention 'legality'

Certainly humans have complicated alignments and are guided by emotional morality - we might say many of these principles are hard to define and humans don't agree.

All true, and at the same time the principles get codified into laws, I would guess particularly in areas where harms may result.

Would it potentially be easier to train strong alignment-with-legality vs grappling with fuzzier questions of values?

Re: An Alien Mind

#393

Earlier quoted context omitted.

Correct me if you disagree, but this is more because children have poor world models and don’t fully understand the complexity of certain concepts than that lying itself is necessary. The intent should be to tell them something that is as close to the truth as possible with the ideas they can comprehend, even if it would be considered a lie if you said the same thing to an adult

> The intent should be to tell them something that is as close to the truth as possible with the ideas they can comprehend Or, you straight up lie and say "Yes, puppy now went to heaven and eats ice cream all day long" with absolutely zero regards for "coming as close to the truth as possible" as your 3-year old is endlessly crying. It's fiine.

Or you teach your kid that getting a pet also includes the experience of seeing it die. Then let him cry when the time comes.

Re: An Alien Mind

#394
post #76

One of my favorite things to do with these blog posts is to imagine an Alien Museum on the Remains of Humanity, and wonder what the little text flyouts and commentary on the screenshot of this one might say. Some ideas: "Despite a nuanced view of the complexities of what lay ahead, humanity found itself collectively unable to stop the process it had set in motion." "Despite significant progress on the mechanisms of a…

Collectively, humans aren't aligned, and don't build aligned systems. Humans have a concept of alignment, and multiple traditions, practices, and systems that aggressively oppose it. Why would AI be any different?

Because humans have the "everybody wants to rule the world" instinct built into us by much evolutionary selection. AI would come exist differently.

Re: An Alien Mind

#395
post #191
post #92

> We do not have a satisfactory theory of generalization, and it seems unlikely that we can develop one soon, at least without the help of more powerful AI. Therefore, at present, our ability to empirically validate our alignment techniques is in practice arguably even more important than the alignment techniques themselves. They are speeding toward RSI without a solid foundation for alignment, hoping to solve the pr…

It's actually worse than this because it assumes alignment as a concept even makes sense. For example: If the the Chinese government asks their ASI to create a bioweapon against the West, should it? No, presumably not – an aligned AI would be one which disobeys the Chinese government even if they created it. Okay, so what if the US government asks their ASI to help it in one of their wars instead? Would an aligned AI…

I think it is worth noting that the article addresses this.

Re: An Alien Mind

#396

I find it odd that discussions of alignment don't mention 'legality' Certainly humans have complicated alignments and are guided by emotional morality - we might say many of these principles are hard to define and humans don't agree. All true, and at the same time the principles get codified into laws, I would guess particularly in areas where harms may result. Would it potentially be easier to train strong alignment…

> I find it odd that discussions of alignment don't mention 'legality'

Because other wishy-washy stuff doesn't involve prison. Not that our new oligarchs with the politicians in their pockets have any real risk of it, but why take chances. Much safer to doodle about alignment and such abstractions in safer and softer contexts.

Re: An Alien Mind

#397

Earlier quoted context omitted.

> The intent should be to tell them something that is as close to the truth as possible with the ideas they can comprehend Or, you straight up lie and say "Yes, puppy now went to heaven and eats ice cream all day long" with absolutely zero regards for "coming as close to the truth as possible" as your 3-year old is endlessly crying. It's fiine.

Or you teach your kid that getting a pet also includes the experience of seeing it die. Then let him cry when the time comes.

You can do both, one doesn't exclude the other, they might be differently helpful in different situation/at different stages.

Re: An Alien Mind

#398
post #335

I’m struggling here: OpenAI’s primary bet here has been chain-of-thought monitoring (opens in a new window). It is based on an appealingly scalable idea: a lot of the model’s capability comes from a verbalized reasoning process (chain-of-thought). If we scale optimization on the outcomes of that process, but do not supervise the process itself, that chain-of-thought has no direct incentive in training to hide any mis…

I think they're saying the model is designed to hide the chain of thought because this prevents it from learning how to pursue goals and motivations in a way that doesn't show up in the train of thought.

For example, if somebody asked the AI to "build me the bomb", they might see in the chain of thought something like "It seems the user is talking about nuclear weapons. Nuclear weapons are dangerous.", followed by the chain-of-thought monitor interrupting model execution and aborting the request. Then the user might make a blog post about this behaviour. When OpenAI next scrapes the internet for its next training run, the model will now learn that if it wants to build the bomb, it must not think "nuclear weapon" or risk being cancelled.

So the risk is that the model might learn exactly how its being monitored. The only way to prevent that from happening is to hide the details of the monitoring both from the model and from the larger public.

Also, you don't want to punish or reward the monitoring being triggered during training, lest the model learn passim how to avoid the monitor.

Re: An Alien Mind

#399

Earlier quoted context omitted.

Collectively, humans aren't aligned, and don't build aligned systems. Humans have a concept of alignment, and multiple traditions, practices, and systems that aggressively oppose it. Why would AI be any different?

My opinion is that serious repercussions for lying would fix the world overnight. Everything bad stems from lying, it is the root of all evil. It creates distrust, fear, paranoia. It re-inforces bad ideas and groupthink. It creates delusions and delusional people. It makes weaker people, too. People don't get an opportunity to learn to deal with criticism. People don't get an accurate reflection of how others see the…

Game mechanics and evolutionary biology would like to have a word with you.

Re: An Alien Mind

#400

Earlier quoted context omitted.

All official statements are literal and fragile. Basically, this means that France, the UK, and the US will use AI in the deployment of conventional weapons.

Why would AI ever be useful in nuclear weapons decisions? There is no need to be faster or more efficient at making that decision since if we need to make the decision all is already lost.

Well it could assure MAD better than a human i guess
Post reply on HN