Live data from Hacker News

Jan Leike joins Anthropic on their superalignment team

twitter.com

11–20 of 39 posts

Re: Jan Leike joins Anthropic on their superalignment team

#11
post #7

Earlier quoted context omitted.

Is there a difference between "superalignment" and "alignment" ?

Yes. “Superalignment” (admittedly a corny term) refers to the specific case of aligning AI systems that are more intelligent than human beings. Alignment is an umbrella term which can also refer to basic work like fine-tuning an LLM to follow instructions.

Huh! All this time I thought the "super" was just for branding/differentiation.

Re: Jan Leike joins Anthropic on their superalignment team

#13

I keep getting Anthropic and Extropic (Guillaume Verdon / Beff Jezos) names mixed up. Anthropic is Claude and Extropic is Thermodynamic hardware many orders of magnitude faster and more energy efficient than CPUs/GPUs.* * parameterized stochastic analog circuits that implement energy-based models (EBMs). Stochastic computing is a computing paradigm that represents numbers using the probability of ones in a bitstream.

> Thermodynamic hardware many orders of magnitude faster and more energy efficient than CPUs/GPUs. I’m sorry, but is this thermodynamic hardware real? Are there any benchmarks? Those claims are pretty strong.

Well you've got Beff Jezos. This is as real as it gets.

Re: Jan Leike joins Anthropic on their superalignment team

#14
post #8

"Automated alignment research" suggests he's still interested in following the superalignment blueprint from OpenAI. So what do you do while you're waiting for the AI that's capable of doing alignment research for you to arrive? If you believe this is a viable path, what's the point of putzing around doing your own research when you'll allegedly have an army of AI researchers at your command in the near future?

I think his tweet can be read as "research in (1) scalable oversight, (2) weak-to-strong generalization, and (3) automated alignment".

Re: Jan Leike joins Anthropic on their superalignment team

#15

I keep getting Anthropic and Extropic (Guillaume Verdon / Beff Jezos) names mixed up. Anthropic is Claude and Extropic is Thermodynamic hardware many orders of magnitude faster and more energy efficient than CPUs/GPUs.* * parameterized stochastic analog circuits that implement energy-based models (EBMs). Stochastic computing is a computing paradigm that represents numbers using the probability of ones in a bitstream.

yes, one is a real company and one is...

Re: Jan Leike joins Anthropic on their superalignment team

#16
post #11
post #7

Earlier quoted context omitted.

Yes. “Superalignment” (admittedly a corny term) refers to the specific case of aligning AI systems that are more intelligent than human beings. Alignment is an umbrella term which can also refer to basic work like fine-tuning an LLM to follow instructions.

Huh! All this time I thought the "super" was just for branding/differentiation.

[deleted]

Re: Jan Leike joins Anthropic on their superalignment team

#17
post #7

Earlier quoted context omitted.

Is there a difference between "superalignment" and "alignment" ?

Yes. “Superalignment” (admittedly a corny term) refers to the specific case of aligning AI systems that are more intelligent than human beings. Alignment is an umbrella term which can also refer to basic work like fine-tuning an LLM to follow instructions.

Is this not something of an oxymoron? If there exists an ai that is more intelligent than humans, how could we mere mortals hope to control it? If we hinder it so that it cannot act in ways that harm humans, can we really be said to have created superintelligence?

It seems to me that the only way to achieve superalignment is to not create superintelligence, if that is even within our control.

Re: Jan Leike joins Anthropic on their superalignment team

#18
post #8

"Automated alignment research" suggests he's still interested in following the superalignment blueprint from OpenAI. So what do you do while you're waiting for the AI that's capable of doing alignment research for you to arrive? If you believe this is a viable path, what's the point of putzing around doing your own research when you'll allegedly have an army of AI researchers at your command in the near future?

> what do you do while you're waiting for the AI that's capable of doing alignment research for you to arrive

Nobody interested in superalignment is interested in waiting until actually threatening AI gets here.

Re: Jan Leike joins Anthropic on their superalignment team

#19
post #17
post #7

Earlier quoted context omitted.

Yes. “Superalignment” (admittedly a corny term) refers to the specific case of aligning AI systems that are more intelligent than human beings. Alignment is an umbrella term which can also refer to basic work like fine-tuning an LLM to follow instructions.

Is this not something of an oxymoron? If there exists an ai that is more intelligent than humans, how could we mere mortals hope to control it? If we hinder it so that it cannot act in ways that harm humans, can we really be said to have created superintelligence? It seems to me that the only way to achieve superalignment is to not create superintelligence, if that is even within our control.

Many people share your views, but others believe it is possible.

Re: Jan Leike joins Anthropic on their superalignment team

#20
post #11
post #7

Earlier quoted context omitted.

Yes. “Superalignment” (admittedly a corny term) refers to the specific case of aligning AI systems that are more intelligent than human beings. Alignment is an umbrella term which can also refer to basic work like fine-tuning an LLM to follow instructions.

Huh! All this time I thought the "super" was just for branding/differentiation.

That was definitely part of it.
Post reply on HN