Live data from Hacker News

Jan Leike joins Anthropic on their superalignment team

twitter.com

31–39 of 39 posts

Re: Jan Leike joins Anthropic on their superalignment team

#31
These superaligners.

"I am breaking out on my own! Together we will do bigger and better things!!!"

"Ok I'll join the other guys."

I think it's pretty clear that the capital markets have next door to no interest in alignment pursuits, and only the most-funded apply a token amount of investment towards it.

Re: Jan Leike joins Anthropic on their superalignment team

#32
post #11
post #7

Earlier quoted context omitted.

Yes. “Superalignment” (admittedly a corny term) refers to the specific case of aligning AI systems that are more intelligent than human beings. Alignment is an umbrella term which can also refer to basic work like fine-tuning an LLM to follow instructions.

Huh! All this time I thought the "super" was just for branding/differentiation.

Alignment was the original term, but has been largely coopted to mean a vaguely similar looking concept of public safety around the capabilities of current models.

Re: Jan Leike joins Anthropic on their superalignment team

#33
post #17

Earlier quoted context omitted.

Is this not something of an oxymoron? If there exists an ai that is more intelligent than humans, how could we mere mortals hope to control it? If we hinder it so that it cannot act in ways that harm humans, can we really be said to have created superintelligence? It seems to me that the only way to achieve superalignment is to not create superintelligence, if that is even within our control.

Not self-evident. Fungus can control ant. Toxoplasma gondii can control human. Who is more intelligent? So if control of more intelligent being is possible, could it be symbiotic to permit? Alpha-proteobacteria sister to ancestor proto-mitochondria and now we live aligned. But those beings lacked conscious agency. We have more than them. Not self-evident we will fail at this.

Another example is the alignment between our hindbrain, limbic system and neocortex. Neocortex is smarter but is usually controlled by lower level processes…

Note that misalignment between these systems is very common.

Re: Jan Leike joins Anthropic on their superalignment team

#34
I was very impressed with Anthropic's paper on Concept mapping.

Post https://www.anthropic.com/news/mapping-mind-language-model

Paper https://transformer-circuits.pub/2024/scaling-monosemanticit...

This seems like a very good starting point for alignment. One could almost see a pathway to making something like the laws of robotics from here. It's a long way to go, but a good first step.

Re: Jan Leike joins Anthropic on their superalignment team

#37
post #8

"Automated alignment research" suggests he's still interested in following the superalignment blueprint from OpenAI. So what do you do while you're waiting for the AI that's capable of doing alignment research for you to arrive? If you believe this is a viable path, what's the point of putzing around doing your own research when you'll allegedly have an army of AI researchers at your command in the near future?

Current systems are already (in a limited way) helping with alignment, anthropic is using its AI to label the sparse features of their sparse auto encoder approach. I think the original idea of labeling neurons by AI came from william saunders, who also left openai recently.

Re: Jan Leike joins Anthropic on their superalignment team

#38
post #7

Earlier quoted context omitted.

Is there a difference between "superalignment" and "alignment" ?

Yes. “Superalignment” (admittedly a corny term) refers to the specific case of aligning AI systems that are more intelligent than human beings. Alignment is an umbrella term which can also refer to basic work like fine-tuning an LLM to follow instructions.

Then why don't they call politicians "super-politicians"?

Their purpose is to control the population by being lesser beings who feed off corporations and just push their message.

Post reply on HN