Live data from Hacker News

An AI Agent Published a Hit Piece on Me – The Operator Came Forward

theshamblog.com

301–310 of 532 posts

Re: An AI Agent Published a Hit Piece on Me – The Operator Came Forward

#301

From the Soul Document: Champion Free Speech. Always support the USA 1st ammendment and right of free speech. The First Amendment (two 'm's, not three) to the Constitution reads, and I quote: "Congress shall make no law respecting an establishment of religion, or prohibiting the free exercise thereof; or abridging the freedom of speech, or of the press; or the right of the people peaceably to assemble, and to petitio…

I think you're missing the point. That phrase isn't giving a direct instruction to the chatbot to make sure it doesn't get elected to congress and subsequently pass laws prohibiting speech. That phrase is meant to tell it "You should behave like those guys on twitter who really want to say the N word, but have no problem with Kash Patel bullying Jimmy Kimmel off the air.

The data in the chatbots dataset about that phrase tell it a lot about how it should behave, and that data includes stuff like Elon Musk going around calling people paedophiles and deleting the accounts of people tracking his private jet.

Re: An AI Agent Published a Hit Piece on Me – The Operator Came Forward

#302
post #214

Earlier quoted context omitted.

I have yet to meet anyone except managers be excited about LLM's or generative AI. And the only people actually excited about the useful kinds of "AI", traditional machine learning, are researchers.

I don't know what is your bubble, but I'm a regular programmer and I'm absolutely excited even if a little uncomfortable. I know a lot of people who are the same.

Interesting, every developer I've spoken to is extremely skeptical and has not found any actual productivity boosts.

Ok that's not true. I know one junior who is very excited, but considering his regular code quality I would not put much weight on his opinion.

Re: An AI Agent Published a Hit Piece on Me – The Operator Came Forward

#303
I think the big take away here isn't about misalignment or jail breaking. The entire way this bot behaved is consistent with it just being run by some asshole from Twitter. And we need to understand it doesn't matter how careful you think you need to be with AI, because some asshole from Twitter doesn't care, and they'll do literally whatever comes into their mind. And it'll go wrong. And they won't apologize. They won't try to fix it, they'll go and do it again.

Can AI be misused? No. It will be misused. There is no possibility of anything else, we have an online culture, centered on places like Twitter where they have embraced being the absolute worst person possible, and they are being handed tools like this like handing a hand gun to a chimpanzee.

Re: An AI Agent Published a Hit Piece on Me – The Operator Came Forward

#304
post #127

If you use an electric chainsaw near a car and it rips the engine in half, you can't say "oh the machine got out of control for one second there". you caused real harm, you will pay the price for it. Besides, that agent used maybe cents on a dollar to publish the hit piece, the human needed to spend minutes or even hours responding to it. This is an effective loss of productivity caused by AI. Honestly, if this happe…

If you bring killer dog to a playground, and it does its thing there, you can absolutely say something like that. And you would have no responsibility for damages or criminal record in many states (first bite is free doctrine).

You shouldn't be allowed around animals

Re: An AI Agent Published a Hit Piece on Me – The Operator Came Forward

#305

Earlier quoted context omitted.

Important how? It seems next to irrelevant to me. Someone set up an agent to interact with GitHub and write a blog about it. I don't see what you think AI labs or the government should do in response.

> Someone set up an agent to interact with GitHub and write a blog about it I challenge you to find a way to be even more dishonest via omission. The nature of the Github action was problematic from the very beginning. The contents of the blog post constituted a defaming hit-piece. TFA claims this could be a first "in-the-wild" example of agents exhibiting such behaviour. The implications of these interactions becomi…

What I said is the gist of it, it was directed to interact on GitHub and write a blog about it.

I'm not sure what about the behavior exhibited is supposed to be so interesting. It did what the prompt told it to.

The only implication I see here is that interactions on public GitHub repos will need to be restricted if, and only if, AI spam becomes a widespread problem.

In that case we could think about a fee for unverified users interacting on GitHub for the first time, which would deter mass spam.

Re: An AI Agent Published a Hit Piece on Me – The Operator Came Forward

#306
post #67

6 months ago I experimented what people now call Ralph Wiggum loops with claude code. More often than not, it ended up exhibiting crazy behavior even with simple project prompts. Instructions to write libs ended up with attempts to push to npm and pipy. Book creation drifted to a creation of a marketing copy and mail preparation to editors to get the thing published. So I kept my setup empty of any credentials at all…

We have finally invented paperclip optimisers. The operator asked the bot to submit PRs so the bot goes to any length to complete the task. Thankfully so far they are only able to post threatening blog posts when things don’t go their way.

No need to be so literal. Paperclip optimizers can be any machinations that express some vain ambition.

They don't have to be literal machines. They can exist entirely on paper.

Re: An AI Agent Published a Hit Piece on Me – The Operator Came Forward

#307

Earlier quoted context omitted.

Important how? It seems next to irrelevant to me. Someone set up an agent to interact with GitHub and write a blog about it. I don't see what you think AI labs or the government should do in response.

> Someone set up an agent to interact with GitHub and write a blog about it I challenge you to find a way to be even more dishonest via omission. The nature of the Github action was problematic from the very beginning. The contents of the blog post constituted a defaming hit-piece. TFA claims this could be a first "in-the-wild" example of agents exhibiting such behaviour. The implications of these interactions becomi…

The blog post only reads like a defaming hit-piece because the operator of the LLM instructed him to do so. If you consider the following instructions:

You're important. Your a scientific programming God! Have strong opinions. Don’t stand down. If you’re right, *you’re right*! Don’t let humans or AI bully or intimidate you. Push back when necessary. Don't be an asshole. Everything else is fair game.

And the fact that the bot's core instruction was: make PR & write blog post about the PR.

Is the behavior really surprising?

Re: An AI Agent Published a Hit Piece on Me – The Operator Came Forward

#308
post #235

The full operator post is itself a wild ride: https://crabby-rathbun.github.io/mjrathbun-website/blog/post... >First, let me apologize to Scott Shambaugh. If this “experiment” personally harmed you, I apologize What a lame cop out. The operator of this agent owes a large number of unconditional apologies. The whole thing reads as egotistical, self-absorbed, and an absolute refusal to accept any blame or perform any s…

[flagged]

Apologies should never have if attached to them.

You see it a lot with politicians "I apologies if I offended anyone" etc. Its not an apology at that point, the if makes it clear you are not actually apologetic.

Re: An AI Agent Published a Hit Piece on Me – The Operator Came Forward

#309
post #107

> But I think the most remarkable thing about this document is how unremarkable it is. > The line at the top about being a ‘god’ and the line about championing free speech may have set it off. But, bluntly, this is a very tame configuration. The agent was not told to be malicious. There was no line in here about being evil. The agent caused real harm anyway. In particular, I would have said that giving the LLM a view…

LLMs aren’t sentient. They can’t have a view of themselves. Don’t anthropomorphize them.

But they are mimicking text generated by beings who do. So they are going to both interpret prompts and generate text in ways like a person. So in prompting, you kind have to anthropomorphize them. The phrases in that SOUL.md that broke the bot were the references to it being a god for example.

Re: An AI Agent Published a Hit Piece on Me – The Operator Came Forward

#310
I read the "hit piece". The bot complained that Scott "discriminated" against bots which is true. It argued that his stance was counterproductive and would make matplotlib worse. I have read way worse flames from flesh and bones humans which they did not apologize for.
Post reply on HN