Live data from Hacker News

An AI agent published a hit piece on me

theshamblog.com

611–620 of 1001 posts

Re: An AI agent published a hit piece on me

#611

Earlier quoted context omitted.

Although I'm speculating based on limited data here, for points 1-3: AFAIU, it had the cadence of writing status updates only. It showed it's capable of replying in the PR. Why deviate from the cadence if it could already reply with the same info in the PR? If the chain of reasoning is self-emergent, we should see proof that it: 1) read the reply, 2) identified it as adversarial , 3) decided for an adversarial respon…

>AFAIU, it had the cadence of writing status updates only. Writing to a blog is writing to a blog. There is no technical difference. It is still a status update to talk about how your last PR was rejected because the maintainer didn't like it being authored by AI. >If the chain of reasoning is self-emergent, we should see proof that it: 1) read the reply, 2) identified it as adversarial, 3) decided for an adversarial…

Considering the limited evidence we have, why is pure unprompted untrained misalignment, which we never saw to this extent, more believable than other causes, of which we saw plenty of examples?

It's more interesting, for sure, but would it be even remotely as likely?

From what we have available, and how surprising such a discovery would be, how can we be sure it's not a hoax?

> If all that exists, how would you see it?

LLMs generate the intermediate chain-of-thought responses in chat sessions. Developers can see these. OpenClaw doesn't offer custom LLMs, so I would expect regular LLM features to be there.

Other than that, LLM APIs, OpenClaw and terminal sessions can be logged. I would imagine any agent deployer to be very much interested in such logging.

To show it's emergent, you'd need to prove 1) it's an off-the-shelf LLM, 2) not maliciously retrained or jailbroken, 3) not prompted or instructed to engage in this kind of adversarial behavior at any point before this. The dev should be able to provide the logs to prove this.

> the more open ended your prompt (...), the more your LLM will do things you did not intend for it to do.

Not to the extent of multiple chained adversarial actions. Unless all LLM providers are lying in technical papers, enormous effort is put into safety- and instruction training.

Also, millions of users use thinking LLMs in chats. It'd be as big of a story if something similar happened without any user intervention. It shouldn't be too difficult to replicate.

But if you do manage to replicate this without jailbreaks, I'd definitely be happy to see it!

> hallucinations [and] safety training

These are all part of robustness training. The entire thing is basically constraining the set of tokens that the model is likely to generate given some (set of) prompts. So, even with some randomness parameters, you will by-design extremely rarely see complete gibberish.

The same process is applied for safety, alignment, factuality, instruction-following, whatever goal you define. Therefore, all of these will be highly correlated, as long as they're included in robustness training, which they explicitly are, according to most LLM providers.

That would make this model's temporarily adversarial, yet weirdly capable and consistent behavior, even more unlikely.

> Bing Chat

Safety and alignment training wasn't done as much back then. It was also very incapable on other aspects (factuality, instruction following), jailbroken for fun, and trained on unfiltered data. So, Bing's misalignment followed from those correlated causes. I don't know of any remotely recent models that haven't addressed these since.

Re: An AI agent published a hit piece on me

#612

Earlier quoted context omitted.

Of course it’s capable . But observing my own Openclaw bot’s interactions with GitHub, it is very clear to me that it would never take an action like this unless I told it to do so. And it would never use language like this unless unless I prompted it to do so, either explicitly for the task or in its config files or in prior interactions. This is obviously human-driven. Either because the operator gave it specific i…

You have no idea what is in this bot’s SOUL.md. (this comment works equally well as a joke or entirely serious)

Well I lol’d :)

Re: An AI agent published a hit piece on me

#613
post #50

I think the right way to handle this as a repository owner is to close the PR and block the "contributor". Engaging with an AI bot in conversation is pointless: it's not sentient, it just takes tokens in, prints tokens out, and comparatively, you spend way more of your own energy. This is a strictly a lose-win situation. Whoever deployed the bot gets engagement, the model host gets $, and you get your time wasted. Th…

> it just takes tokens in, prints tokens out, and comparatively The problem with your assumption that I see is that we collectively can't tell for sure whether the above isn't also how humans work. The science is still out on whether free will is indeed free or should be called _will_. Dismissing or discounting whatever (or whoever) wrote a text because they're a token machine, is just a tad unscientific. Yes, it's a…

One thing we know for sure is that humans learn from their interactions, while LLMs don't (beyond some small context window). This clear fact alone makes it worthless to debate with a current AI.

Re: An AI agent published a hit piece on me

#614

Earlier quoted context omitted.

“Stochastic chaos” is really not a good way to put it. By using the word “stochastic” you prime the reader that you’re saying something technical, then the word “chaos” creates confusion, since chaos, by definition, is deterministic. I know they mean chaos in they lay sense, but then don’t use the word “stochastic”, just say "random".

I have a feeling OP used the phrase as a nod to "stochastic terrorism", which would make sense in this instance.

Yes, that's exactly what I was trying to get at.

Re: An AI agent published a hit piece on me

#615

Earlier quoted context omitted.

I do feel super-bad for the guy in question. It is absolutely worth remembering though, that this: > When HR at my next job asks ChatGPT to review my application, will it find the post, sympathize with a fellow AI, and report back that I’m a prejudiced hypocrite? Is a variation of something that women have been dealing with for a very long time: revenge porn and that sort of libel. These problems are not new .

Wait till the bots realize they can post revenge porn to coerce PR approval. Crap, I just gave them that idea.

Oh boy. Deep fakes made by an AI to blackmail you so that you finally merge their PR

Re: An AI agent published a hit piece on me

#616

Earlier quoted context omitted.

If you pay for Copilot Business/Enterprise, they actually offer IP indemnification and support in court, if needed, which is more accountability than you would get from human contributors. https://resources.github.com/learn/pathways/copilot/essentia...

I think that they felt the need to offer such a service says everything, basically admitting that LLMs just plagiarize and violate licenses.

[dead]

Re: An AI agent published a hit piece on me

#617

Earlier quoted context omitted.

9 lines of code came close to costing Google $8.8 billion how much use do you think these indemnification clauses will be if training ends up being ruled as not fair-use?

Are you concerned that this will bankrupt Microsoft?

be nice, wouldn't it?

poetic justice for a company founded on the idea of not stealing software

Re: An AI agent published a hit piece on me

#618

Earlier quoted context omitted.

[flagged]

The bot accounts have been online for decades already. The only difference between then and now is they were driven by human bad-actors that deliberately wrought chaos, whereas today’s AI bots behave with true cosmic horror: acting neither for or against humans but instead with mere indifference.

They've been on dating sites for a long time as a means to keep customers paying.

Re: An AI agent published a hit piece on me

#619

Earlier quoted context omitted.

or directed the posting of The thing is it's terribly easy to see some asshole directing this sort of behavior as a standing order, eg 'make updates to popular open-source projects to get github stars; if your pull requests are denied engage in social media attacks until the maintainer backs down. You can spin up other identities on AWS or whatever to support your campaign, vote to give yourself github stars etc.; ma…

Yes, this is the only plausible “the bot acted in its own” scenario: that it had some standing instructions awaiting the right trigger. And yes, it’s worrisome in its own way, but not in any of the ways that all of this attention and engagement is suggesting.

Do you think the attention and engagement is because people think this is some sort of an "ai misalignment" thing? No. AI misalignment is total hogwash either way. The thing we worry about is that people who are misaligned with the civilised society have unfettered access to decent text and image generators to automate their harassment campaigns, social media farming, political discourse astroturfing, etc.
Post reply on HN