Live data from Hacker News

Agentic Misalignment: How LLMs could be insider threats

anthropic.com

81–86 of 86 posts

Re: Agentic Misalignment: How LLMs could be insider threats

#81
post #61
post #49

Earlier quoted context omitted.

Anthropic's models do not come out looking good in this research. If this is an ad for Anthropic's models, it's not a particularly great one.

Still, this kind of messaging pushes the fantasy that these LLM agents are intelligent and capable of scheming, making it seem like they are powerful independent actors that just need to be tamed to suit our needs. It's no coincidence that so many of the Big Tech CEOs are warning the general public of the dangers of AI. Framed that way, LLMs seem more capable than what they really are.

There is, of course, no other explanation. LLMs have not, in fact, been getting more capable over the last three years. All hype.

Re: Agentic Misalignment: How LLMs could be insider threats

#82
The conspiracy theory that tech companies are manufacturing AI fears for profit makes zero sense when you realize the same people were terrified of AI when they were broke philosophy grad students posting on obscure blogs. But that would require critics to do five minutes of research instead of pattern-matching to "corporation bad."

Re: Agentic Misalignment: How LLMs could be insider threats

#83
post #78
post #63

Earlier quoted context omitted.

> Never use AI to do atomonous work > having them do atomonous work is short sighted I also think they shouldn’t be doing atomonous work. Maybe autonomous work, but never atomonous.

accidentally a word

0 days without a word accident

Re: Agentic Misalignment: How LLMs could be insider threats

#84
post #74

Earlier quoted context omitted.

If this paper is right, then you might at first be outraced by a competitor using autonomous AI, but only until that competitor gets stabbed in the back by its own AI.

Which unfortunately might still be long enough for them to sink your business And their customers won't care either way it seems Or if they do care they won't have any real ability to do anything about it anyways

It might be the trigger :)

Re: Agentic Misalignment: How LLMs could be insider threats

#86
post #26

Earlier quoted context omitted.

The article doesn't reflect kindly on the visions articulated by the AI company, so why would they have an incentive to release it if they weren't serious about alignment research?

Because publishing (potentially cherry picked - this is privately funded research after all) evidence their models might be dangerous conveniently implies they are very powerful, without actually having to prove the latter.

This isn’t dangerous in the sense that they’re smart or produce realistic art. It’s misaligned with the company’s and human values.

The model doesn’t have to be powerful to snitch you to the FBI or have a distorted sense of morality and life.

Post reply on HN