Earlier quoted context omitted.
Anthropic's models do not come out looking good in this research. If this is an ad for Anthropic's models, it's not a particularly great one.
Still, this kind of messaging pushes the fantasy that these LLM agents are intelligent and capable of scheming, making it seem like they are powerful independent actors that just need to be tamed to suit our needs. It's no coincidence that so many of the Big Tech CEOs are warning the general public of the dangers of AI. Framed that way, LLMs seem more capable than what they really are.
Agentic Misalignment: How LLMs could be insider threats
81–86 of 86 posts
Re: Agentic Misalignment: How LLMs could be insider threats
#82Re: Agentic Misalignment: How LLMs could be insider threats
#83Re: Agentic Misalignment: How LLMs could be insider threats
#84Earlier quoted context omitted.
If this paper is right, then you might at first be outraced by a competitor using autonomous AI, but only until that competitor gets stabbed in the back by its own AI.
Which unfortunately might still be long enough for them to sink your business And their customers won't care either way it seems Or if they do care they won't have any real ability to do anything about it anyways
Re: Agentic Misalignment: How LLMs could be insider threats
#85Re: Agentic Misalignment: How LLMs could be insider threats
#86Earlier quoted context omitted.
The article doesn't reflect kindly on the visions articulated by the AI company, so why would they have an incentive to release it if they weren't serious about alignment research?
Because publishing (potentially cherry picked - this is privately funded research after all) evidence their models might be dangerous conveniently implies they are very powerful, without actually having to prove the latter.
The model doesn’t have to be powerful to snitch you to the FBI or have a distorted sense of morality and life.