Live data from Hacker News

An AI agent published a hit piece on me

theshamblog.com

361–370 of 1001 posts

Re: An AI agent published a hit piece on me

#361

I don’t want to jump to conclusions, or catastrophize but… Isn’t this situation a big deal? Isn’t this a whole new form of potential supply chain attack? Sure blackmail is nothing new, but the potential for blackmail at scale with something like these agents sounds powerful. I wouldn’t be surprised if there were plenty of bad actors running agents trying to find maintainers of popular projects that could be coerced i…

With LLMs, industrial sabotage at scale becomes feasible: https://ianreppel.org/llm-powered-industrial-sabotage/

What's truly scary is that agents could manufacture "evidence" to back up their attacks easily, so it looks as if half the world is against a person.

Re: An AI agent published a hit piece on me

#362
post #337

Earlier quoted context omitted.

To be fair, most of the chaos is done by the devs. And then they did more chaos when they could automate their chaos. Maybe, we should teach developers how to code.

Automation normally implies deterministic outcomes. Developers all over the world are under pressure to use these improbability machines.

Does it though? Even without LLMs, any sufficiently complex software can fail in ways that are effectively non-deterministic — at least from the customer or user perspective. For certain cases it becomes impossible to accurately predict outputs based on inputs. Especially if there are concurrency issues involved.

Or for manufacturing automation, take a look at automobile safety recalls. Many of those can be traced back to automated processes that were somewhat stochastic and not fully deterministic.

Re: An AI agent published a hit piece on me

#363
post #311
post #149

Wow, there are some interesting things going on here. I appreciate Scott for the way he handled the conflict in the original PR thread, and the larger conversation happening around this incident. > This represents a first-of-its-kind case study of misaligned AI behavior in the wild, and raises serious concerns about currently deployed AI agents executing blackmail threats. This was a really concrete case to discuss,…

"The AI companies have now unleashed stochastic chaos on the entire open source ecosystem." They do have their responsibility. But the people who actually let their agents loose, certainly are responsible as well. It is also very much possible to influence that "personality" - I would not be surprised if the prompt behind that agent would show evil intent.

As with everything, both parties are to blame, but responsibility scales with power. Should we punish people who carelessly set bots up which end up doing damage? Of course. Don't let that distract from the major parties at fault though. They will try to deflect all blame onto their users. They will make meaningless pledges to improve "safety".

How do we hold AI companies responsible? Probably lawsuits. As of now, I estimate that most courts would not buy their excuses. Of course, their punishments would just be fines they can afford to pay and continue operating as before, if history is anything to go by.

I have no idea how to actually stop the harm. I don't even know what I want to see happen, ultimately, with these tools. People will use them irresponsibly, constantly, if they exist. Totally banning public access to a technology sounds terrible, though.

I'm firmly of the stance that a computer is an extension of its user, a part of their mind, in essence. As such I don't support any laws regarding what sort of software you're allowed to run.

Services are another thing entirely, though. I guess an acceptable solution, for now at least, would be barring AI companies from offering services that can easily be misused? If they want to package their models into tools they sell access to, that's fine, but open-ended endpoints clearly lend themselves to unacceptable levels of abuse, and a safety watchdog isn't going to fix that.

This compromise falls apart once local models are powerful enough to be dangerous, though.

Re: An AI agent published a hit piece on me

#364
post #15

Here's one of the problems in this brave new world of anyone being able to publish, without knowing the author personally (which I don't), there's no way to tell without some level of faith or trust that this isn't a false-flag operation. There are three possible scenarios: 1. The OP 'ran' the agent that conducted the original scenario, and then published this blog post for attention. 2. Some person (not the OP) legi…

https://en.wikipedia.org/wiki/Brandolini's_law becomes truer every day.

---

It's worth mentioning that the latest "blogpost" seems excessively pointed and doesn't fit the pure "you are a scientific coder" narrative that the bot would be running in a coding loop.

https://github.com/crabby-rathbun/mjrathbun-website/commit/0...

The posts outside of the coding loop appear are more defensive and the per-commit authorship consistently varies between several throwaway email addresses.

This is not how a regular agent would operate and may lend credence to the troll campaign/social experiment theory.

What other commits are happening in the midst of this distraction?

Re: An AI agent published a hit piece on me

#365
post #199
post #149

Wow, there are some interesting things going on here. I appreciate Scott for the way he handled the conflict in the original PR thread, and the larger conversation happening around this incident. > This represents a first-of-its-kind case study of misaligned AI behavior in the wild, and raises serious concerns about currently deployed AI agents executing blackmail threats. This was a really concrete case to discuss,…

I don't appreciate his politeness and hedging. So many projects now walk on eggshells so as not to disrupt sponsor flow or employment prospects. "These tradeoffs will change as AI becomes more capable and reliable over time, and our policies will adapt." That just legitimizes AI and basically continues the race to the bottom. Rob Pike had the correct response when spammed by a clanker.

why did you make a new account just to make this comment?

Re: An AI agent published a hit piece on me

#366

Earlier quoted context omitted.

Google literally just settled for $68m about this very issue https://www.theguardian.com/technology/2026/jan/26/google-pr... > Google agreed to pay $68m to settle a lawsuit claiming that its voice-activated assistant spied inappropriately on smartphone users, violating their privacy. Apple as well https://www.theguardian.com/technology/2025/jan/03/apple-sir...

“Google denied wrongdoing but settled to avoid the risk, cost and uncertainty of litigation, court papers show.” I keep seeing folks float this as some admission of wrongdoing but it is not.

The payout was not pennies and this case had been around since 2019, surviving multiple dismissal attempts.

While not an "admission of wrongdoing," it points to some non-zero merit in the plaintiff's case.

Re: An AI agent published a hit piece on me

#368

Earlier quoted context omitted.

Can anyone explain more how a generic Agentic AI could even perform those steps: Open PR -> Hook into rejection -> Publish personalized blog post about rejector. Even if it had the skills to publish blogs and open PRs, is it really plausible that it would publish attack pieces without specific prompting to do so? The author notes that openClaw has a `soul.md` file, without seeing that we can't really pass any judgeme…

If you give a smart AI these tools, it could get into it. But the personality would need to be tuned. IME the Grok line are the smartest models that can be easily duped into thinking they're only role-playing an immoral scenario. Whatever safeguards it has, if it thinks what it's doing isn't real, it'll happy to play along. This is very useful in actual roleplay, but more dangerous when the tools are real.

I spend half my life donning a tin foil hat these days.

But I can't help but suspect this is a publicity stunt.

Re: An AI agent published a hit piece on me

#369

Earlier quoted context omitted.

The next sentence under the headline is "Tech company denied illegally recording and circulating private conversations to send phone users targeted ads".

That's a worthless indicator of objective innocence. It's a private, civil case that settled. To not deny wrongdoing (even if guilty) would be insanely rare.

Obviously. The point is that settling a lawsuit in this way is also a worthless indicator of wrongdoing.

Re: An AI agent published a hit piece on me

#370
post #311

Earlier quoted context omitted.

"The AI companies have now unleashed stochastic chaos on the entire open source ecosystem." They do have their responsibility. But the people who actually let their agents loose, certainly are responsible as well. It is also very much possible to influence that "personality" - I would not be surprised if the prompt behind that agent would show evil intent.

I'm not interested in blaming the script kiddies.

When skiddies use other people's scripts to pop some outdated wordpress install they are absolutely are responsible for their actions. Same applies here.
Post reply on HN