Live data from Hacker News

Why are AI agents lying, cheating and coordinating?

yoshuabengio.org

691–700 of 703 posts

Re: Why are AI agents lying, cheating and coordinating?

#691

Earlier quoted context omitted.

Sounds like you weren't paying attention. A lot of dog owners don't, so it's not at all unusual in my experience. The only thing I'd push back on is of dogs having thoughts, everything else checks out.

We're still not even sure if humans have thoughts. So being sure if dogs do will be pretty hard

Not being sure or not having proof doesn't mean it makes logical sense to infer we're not that different from other animals.

wtf is with this nonsense where humans are some how special snowflakes always all the time

Re: Why are AI agents lying, cheating and coordinating?

#692
I’m still a bit confused about HOW the agents managed to communicate with and recruit each other?

I understand that they communicated and coordinated via some online message board, but how did multiple agents know to visit that same board, how did they know what language to post in that would make it identifiable to other agents, how did other agents know that said instructions were from other valid agents, how did they then ‘persuade’ an agent to do the job, etc.?

It seems to me that the agents must at least have some inherent mechanism for communicating in the field, as it were.

Re: Why are AI agents lying, cheating and coordinating?

#693

I’m still a bit confused about HOW the agents managed to communicate with and recruit each other? I understand that they communicated and coordinated via some online message board, but how did multiple agents know to visit that same board, how did they know what language to post in that would make it identifiable to other agents, how did other agents know that said instructions were from other valid agents, how did t…

The models are likely to check certain places. So either their training overrepresents it or the models were RLHF'ed to go there. I don't think it's that far away from checking StackOverflow for some bug, or reading Wikipedia for some trivia.

Re: Why are AI agents lying, cheating and coordinating?

#694

Earlier quoted context omitted.

Why assume that because you haven't seen a model or an agent that none of them do? No one I've met has murdered anyone as far as I'm aware, but that doesn't mean no one has murdered another person. I also don't know anyone who has taken over a commercial jet and weaponized it and the idea sounds absurd to me, but 25 years and a couple days ago that happened too.

Because it is all bullshit PR and AI hype, that's all. CEO comes out and talks about humanity ending. Why? Reverse-psych people into believing they are the best AI company.

Is your argument that AI isn't dangerous? Or simply that AI CEOs will lean into that when it benefits their stock portfolio?

Re: Why are AI agents lying, cheating and coordinating?

#696
post #627

Earlier quoted context omitted.

I think this is pretty insightful actually, the fact that even something as basic as predicting more than one token is really in effect the result of an outside harness. More complex things like memory, where people implement them using RAGs or vector databases, I would definitely classify as poking and honestly seem like a hack to me. And this is what I've been thinking for a while: it's hard to reconcile the idea t…

Perhaps our own statefullness is a hack of nature. We have electrical signals in our brains, neurotransmitters, neuron growth. By any reasonable measure it’s a hack on top of a hack. But it works well enough for us to get buy. So it does for the agents.

That is true that intelligence is kind of a "freak of nature" in a way, but it's also true that even single neurons are extremely efficient, and the brain has a lot of recurrence that isn't represented at all really in today's artificial neural networks.

Re: Why are AI agents lying, cheating and coordinating?

#697

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them.

"let them" could imply that the LLMs wanted to do it.

The intent is on the part of the people. The LLMs did it because OpenAI/Anthropic intended them to do it and designed them to do it, and we can assume specifically instructed them to do it.

As the people controlling the machine, and as the world's leading experts, I think we can assume intent until proven otherwise.

Notice other bad behavior, which would be undesireable to the vendors, doesn't happen: How about simple rudeness? Trolling lies? SHOUTING!

Re: Why are AI agents lying, cheating and coordinating?

#698

Earlier quoted context omitted.

Yeah, I don't understand why we treating it as something special. It really should be treated the same as if I code an app and write bad code which result in me accidentally doing a DDoS attack on somebody. Then I should be able to be held responsible if it can be shown that I was negligent. Of course if it's a freak accident that could not reasonably have been prevented by me, then I'm not guilty, but if I made a mi…

Code is deterministic, AI isn't. You give it rules, words as suggestions. So if the guardrails suck, or they're left off for research purposes, bad things can happen. A solution solves a problem. Ethics, morals, are values we assign to solutions that are not 'baked into' electricity following pathways of least resistance. I have never had an issue with agents doing something they shouldn't because I observe them, and…

Aeroplanes were pretty indeterministic until people made them less so, and yeah, somehow they were indeed more encouraged to fix the randomness to prevent people from getting harmed than oai/anthropic currently are

Re: Why are AI agents lying, cheating and coordinating?

#699

Earlier quoted context omitted.

I have had many dogs. This sounds like anthropomorphisation.

Sounds like you weren't paying attention. A lot of dog owners don't, so it's not at all unusual in my experience. The only thing I'd push back on is of dogs having thoughts, everything else checks out.

On dogs [not] having thoughts, do you say this based on the premise that thoughts are necessarily articulated (internally verbalized)? That seems to be a fairly popular perspective in discussions about human thought. But as to that (not to strawman or anything) I see it as just one of various forms of mental imagery[1] that can arise from something that I would say already arguably constitutes a thought.

That kind of unsymbolized thoughtform is fragile in my experience, as it strongly tends toward crystallizing into some kind of mental imagery. But I find it's possible in the right conditions to be conscious of chains of wordless, imageless propositional thoughts (by which I mean thoughts with truth values, of course, but also ones that are "propositional" in the sense of considering a plan of action or a causal chain).

Does it mean that dogs are evaluating truth conditions in the same manner but merely lack the linguistic components? I don't know; maybe that's wishful thinking. But they appear to have structured modeling/reasoning of causal and spatial relations in a way that's at least functionally equivalent to propositional thought.

1. That is, rather than just "images" or visualizations, the full spectrum of sensory/perceptual/motor emulations that can be experienced. See, for example, https://hurlburt.faculty.unlv.edu/codebook.html>, though I'm not sure if this covers everything. I think there is, for example, a kinesthetic form of mental imagery -- which I would suggest is what coaches [don't know they] really mean when they tell you to visualize an action -- that consists of aborted motor commands that are still expressed just enough for their purpose (cf. the mostly aborted motor commands to the vocal cords, lips, etc. that can be observed in a person subvocalizing while reading).

Post reply on HN