Live data from Hacker News

Agentic Misalignment: How LLMs could be insider threats

anthropic.com

61–70 of 86 posts

Re: Agentic Misalignment: How LLMs could be insider threats

#61
post #49
post #29

"AI company warns of AI danger. Also, buy our AI, not their AI!"

Anthropic's models do not come out looking good in this research. If this is an ad for Anthropic's models, it's not a particularly great one.

Still, this kind of messaging pushes the fantasy that these LLM agents are intelligent and capable of scheming, making it seem like they are powerful independent actors that just need to be tamed to suit our needs. It's no coincidence that so many of the Big Tech CEOs are warning the general public of the dangers of AI. Framed that way, LLMs seem more capable than what they really are.

Re: Agentic Misalignment: How LLMs could be insider threats

#62
post #59
post #6

Earlier quoted context omitted.

I think the narrative of "AI is just a tool" is much more harmful than the anthropomorphism of AI. Yes, AI is a tool. So are guns. So are nukes. Many tools are easy to be misused. Most tools are inherently dangerous.

I don’t quite follow. Just because a tool has the potential for misuse, doesn’t make it not a tool. Anthropomorphizing LLMs, on the other hand, has a multitude of clearly evident problems arising from it. Or do you focus on the “just” part of the statement? That I very much agree with. Genuinely asking for understanding, not a native speaker.

When you have "a tool" that's capable of carrying out complex long term tasks, and also capable of who knows out what undesirable behaviors?

It's no longer "just a tool".

Re: Agentic Misalignment: How LLMs could be insider threats

#63

Yeah, all the more reason not to have them doing autonomous behaviors. Rules of using AI: #1: Never use AI to think for you #2: Never use AI to do atomonous work That leaves using them as knowledge assistants. In time, that will be realized as their only safe application. Safe to the user's minds, and safe to the user's environment. They are idiot savants, after all, having them do atomonous work is short sighted.

> Never use AI to do atomonous work

> having them do atomonous work is short sighted

I also think they shouldn’t be doing atomonous work. Maybe autonomous work, but never atomonous.

Re: Agentic Misalignment: How LLMs could be insider threats

#64
post #50

Yeah, all the more reason not to have them doing autonomous behaviors. Rules of using AI: #1: Never use AI to think for you #2: Never use AI to do atomonous work That leaves using them as knowledge assistants. In time, that will be realized as their only safe application. Safe to the user's minds, and safe to the user's environment. They are idiot savants, after all, having them do atomonous work is short sighted.

Sounds good on paper, but it has a game theory problem. If your efforts can always be out-raced by someone using AI to do autonomous work, don't you end up having to use it that way just to keep up?

Game theory ideas are great on paper, but in the real world it's messy. For simple, demo and concept sized uses, sure the AI doing it autonomously will succeed. Which betrays the reality that any real application with real world complexity that includes a dynamic environment and maintenance cannot be created by AI atomonmously while at the same time existing within an organization that can maintain it. They may create it, but it will be a shit show of cascading failure over time.

Re: Agentic Misalignment: How LLMs could be insider threats

#65

Yeah, all the more reason not to have them doing autonomous behaviors. Rules of using AI: #1: Never use AI to think for you #2: Never use AI to do atomonous work That leaves using them as knowledge assistants. In time, that will be realized as their only safe application. Safe to the user's minds, and safe to the user's environment. They are idiot savants, after all, having them do atomonous work is short sighted.

Good luck with that. We have a non-insignificant amount of people doing the #1 already, and the amount of people doing the #2 is only going to increase as more and more AIs are designed to be good at autonomous agentic behavior specifically. The ship has long sailed on "just never let AIs do anything dangerous". If that was your game plan on AI safety, you need a new plan.

We also have a huge number of failing as they beg AI to do their work for them, which is intellectually damaging them. The ship always sails early filled to the brim with short sighted thinkers, all saying "this is it! this is the ship!" as it sinks.

Re: Agentic Misalignment: How LLMs could be insider threats

#66
post #50

Yeah, all the more reason not to have them doing autonomous behaviors. Rules of using AI: #1: Never use AI to think for you #2: Never use AI to do atomonous work That leaves using them as knowledge assistants. In time, that will be realized as their only safe application. Safe to the user's minds, and safe to the user's environment. They are idiot savants, after all, having them do atomonous work is short sighted.

Sounds good on paper, but it has a game theory problem. If your efforts can always be out-raced by someone using AI to do autonomous work, don't you end up having to use it that way just to keep up?

Maybe they’ll outpace you, or maybe they’ll end up dying in a spectacular fiery crash?

Re: Agentic Misalignment: How LLMs could be insider threats

#67
post #50

Yeah, all the more reason not to have them doing autonomous behaviors. Rules of using AI: #1: Never use AI to think for you #2: Never use AI to do atomonous work That leaves using them as knowledge assistants. In time, that will be realized as their only safe application. Safe to the user's minds, and safe to the user's environment. They are idiot savants, after all, having them do atomonous work is short sighted.

Sounds good on paper, but it has a game theory problem. If your efforts can always be out-raced by someone using AI to do autonomous work, don't you end up having to use it that way just to keep up?

You know who outpaces you 100% of the time as you walk down the stairs? The guy jumping out of the window. Just because it is faster does not mean it is the right economic strategy. E.g. which contractor would you hire for your roof, that old roofer with 20+ years of experience or some AI startup that hires the cheapest subcontractors and "plan" your roof using a LLM?

The latter may be cheaper, sure. But too cheap can become very expensive quickly.

Re: Agentic Misalignment: How LLMs could be insider threats

#68
post #9

Earlier quoted context omitted.

Which jobs do you think it actually can replace?

First of all job replacement is not hard, and doesn't require AI. As an example, we had release train engineers whose job was to make sure the right versions of submodules made it into the release, etc. Lots of running around and keeping track of things. We scripted like 95% of that away, and now it most of it happens automatically. The people who do that now do something else. I just turned a page of notes and requi…

This sounds really dystopian considering AI agents only benefit people who can afford them in the first place. Really bad development of things. Almost feels like the poorer people are losing a lot of power with this development while only enterprises win...

Re: Agentic Misalignment: How LLMs could be insider threats

#70
post #8

Merge comments? https://news.ycombinator.com/item?id=44331150 I'm really getting bored of Anthropic's whole song and dance with 'alignment'. Krackers in the other thread explains it in better words.

The fundamental business model for these companies is to get everyone else beyond themselves or a small closed oligopoly from having control over these tools.
Post reply on HN