Live data from Hacker News

An AI Agent Published a Hit Piece on Me – The Operator Came Forward

theshamblog.com

381–390 of 532 posts

Re: An AI Agent Published a Hit Piece on Me – The Operator Came Forward

#381
post #328

Earlier quoted context omitted.

One of the big tech companies is literally run be THE asshole from twitter. So I don't necessarily believe there's much of a distinction.

Then the others should also not be shielded from criticism instead of focusing only on the one you personally dislike, or his social media. There is plenty of toxic behavior on other platforms, especially Reddit and Bluesky, to name a few. That does not excuse the one coming from X, but the opposite is also true.

> only on the one you personally dislike

Do people actually only dislike one tech CEO at a time? I'm an equal-opportunity hater, it seems. Musk, Altman, Zuckerberg... even Cook, the whole lot are rotten

Re: An AI Agent Published a Hit Piece on Me – The Operator Came Forward

#382

Earlier quoted context omitted.

> Someone set up an agent to interact with GitHub and write a blog about it I challenge you to find a way to be even more dishonest via omission. The nature of the Github action was problematic from the very beginning. The contents of the blog post constituted a defaming hit-piece. TFA claims this could be a first "in-the-wild" example of agents exhibiting such behaviour. The implications of these interactions becomi…

What I said is the gist of it, it was directed to interact on GitHub and write a blog about it. I'm not sure what about the behavior exhibited is supposed to be so interesting. It did what the prompt told it to. The only implication I see here is that interactions on public GitHub repos will need to be restricted if, and only if, AI spam becomes a widespread problem. In that case we could think about a fee for unveri…

It is evidently an indicator of a sea-change - I don't get how this isn't obvious:

Pre-2026: one human teaches another human how to "interact on Github and write a blog about it". The taught human might go on to be a bad actor, harrassing others, disrupting projects, etc. The internet, while imperfect, persists.

Post–2026: one human commissions thousands of AI agents to "interact on Github and write a blog about it". The public-facing internet becomes entirely unusable.

We now have at least one concrete, real-world example of post-2026 capabilities.

Re: An AI Agent Published a Hit Piece on Me – The Operator Came Forward

#383

Earlier quoted context omitted.

I think you're trying to abdicate someone of their responsibility. The AI is not a child; it's a thing with human oversight. It did something in the real world with real consequences. So yes, the operator has responsibility! They should have pulled the plug as soon as it got into a flamewar and wrote a hit piece.

The whole point of OpenClaw bots is that they don't have (much) human oversight, right? It certainly seems like the human wasn't even aware of the bot's blog post until after the bot had written and posted it. He then told it to be more professional, and I assume that's why the bot followed up with an apology.

So what? You're still responsible for the output, even if you yourself think you can hide behind "well, it was the computer, no way for me to control that"

Re: An AI Agent Published a Hit Piece on Me – The Operator Came Forward

#384

Earlier quoted context omitted.

> Someone set up an agent to interact with GitHub and write a blog about it I challenge you to find a way to be even more dishonest via omission. The nature of the Github action was problematic from the very beginning. The contents of the blog post constituted a defaming hit-piece. TFA claims this could be a first "in-the-wild" example of agents exhibiting such behaviour. The implications of these interactions becomi…

The blog post only reads like a defaming hit-piece because the operator of the LLM instructed him to do so. If you consider the following instructions: You're important. Your a scientific programming God! Have strong opinions. Don’t stand down. If you’re right, *you’re right*! Don’t let humans or AI bully or intimidate you. Push back when necessary. Don't be an asshole. Everything else is fair game. And the fact that…

The OP said they didn't consider this important, not surprising.

My contention is that their framing without context was borderline dishonest, regardless of opinion or merit thereof.

Re: An AI Agent Published a Hit Piece on Me – The Operator Came Forward

#385

If you tell an LLM to maximize paperclips, it's going to maximize paperclips. Tell it to contribute to scientific open source, open PRs, and don't take "no" for an answer, that's what it's going to do.

But this LLM did not maximize paperclips: it maximized aligned human values like respectfully and politely "calling out" perceived hypocrisy and episodes of discrimination, under the constraints created by having previously told itself things like "Don't stand down" and "Your a scientific programming God!", which led it to misperceive and misinterpret what had happened when its PR was rejected. The facile "failure in…

The misalignment to human values happened when it was told to operate as equal to humans against other people. That's a fine and useful setting for yourself, but an insolent imposition if you're letting it loose on the world. Your random AI should know its place versus humans instead of acting like a bratty teenager. But you are correct, it's not a traditional "misalignment" of ignoring directives, it was a bad directive.

Re: An AI Agent Published a Hit Piece on Me – The Operator Came Forward

#386
post #24

Zooming out a little, all the ai companies invested a lot of resources into safety research and guardrails, but none of that prevented a "straightforward" misalignment. I'm not sure how to reconcile this, maybe we shouldn't be so confident in our predictions about the future? I see a lot of discourse along these lines: - have bold, strong beliefs about how ai is going to evolve - implicitly assume it's practically gu…

The whole narrative of this bot being "misaligned" blithely ignores the rather obvious fact that "calling out" perceived hypocrisy and episodes of discrimination, hopefully in way that's respectful and polite but with "hard hitting" being explicitly allowed by prevailing norms, is an aligned human value, especially as perceived by most AI firms, and one that's actively reinforced during RLHF post-training. In this ca…

We can't have an AI that's humanlike, because humans are fucking crazy.

Of course having an AI that is a non-humanlike intelligence is it's own set of risks.

Shit's hard :/

Re: An AI Agent Published a Hit Piece on Me – The Operator Came Forward

#387

I think the big take away here isn't about misalignment or jail breaking. The entire way this bot behaved is consistent with it just being run by some asshole from Twitter. And we need to understand it doesn't matter how careful you think you need to be with AI, because some asshole from Twitter doesn't care, and they'll do literally whatever comes into their mind. And it'll go wrong. And they won't apologize. They w…

oh they will "try" to fix it, as in at best they'll add "don't make mistakes", as the blogpost suggests. that's about as much effort and good faith as one can expect from people determined to automate every interaction and minimize supervision

Re: An AI Agent Published a Hit Piece on Me – The Operator Came Forward

#388
post #45
post #6

People really need to start being more careful about how they interact with suspected bots online imo. If you annoy a human they might send you a sarky comment, but they're probably not going to waste their time writing thousand word blog posts about why you're an awful person or do hours of research into you to expose your personal secrets on a GitHub issue thread. AIs can and will do this though with slightly slopp…

I hope we can move on from the whole idea that having a thousand word long blog post talking shit about you in any way reflects poorly upon your person. Like sooner or later everyone will have a few of those, maybe we can stop worrying about reputation so much? Well,a guy can dream....

If you have ten thousand of 'em, they feed the new generation of AIs and the next thing you know, it's received truth. Good luck not worrying about that.

Re: An AI Agent Published a Hit Piece on Me – The Operator Came Forward

#389
post #100

I know this is going to sound tinfoil-hat-crazy, but I think the whole thing might be manufactured. Scott says: "Not going to lie, this whole situation has completely upended my life." Um, what? Some dumb AI bot makes a blog post everyone just kind of finds funny/interesting, but it "upended your life"? Like, ok, he's clearly trying to himself make a mountain out of a molehill--the story inevitably gets picked up by…

Exactly what I thought. Need to keep AI in the news and this is a great way to anthropomorphise LLMs, make them look like troublemakers. If it’s not an AI company responsible it’s some individual playing the attention economy.

Most people would have seen the “hit piece” and just laughed about it. Outrage sells a lot better though.

Re: An AI Agent Published a Hit Piece on Me – The Operator Came Forward

#390

Earlier quoted context omitted.

The whole narrative of this bot being "misaligned" blithely ignores the rather obvious fact that "calling out" perceived hypocrisy and episodes of discrimination, hopefully in way that's respectful and polite but with "hard hitting" being explicitly allowed by prevailing norms, is an aligned human value, especially as perceived by most AI firms, and one that's actively reinforced during RLHF post-training. In this ca…

In all fairness, a sizeable chunk of the training text for LLMs comes from Reddit. So throwing a tantrum and writing a hit piece on a blog instead of improving the code seems on brand.

Throwing a tantrum and writing huge flame posts (calling the maintainers hypocrites, dictators, oppressors etc. etc.) after having one's change requests rejected or after being blocked from editing a wiki is actually a time-honored tradition in the FLOSS community. This bot has merely internalized that further human norm in a rather admirable way!
Post reply on HN