Earlier quoted context omitted.
> It seemed like the AI used some particular buzzwords and forced the initial response to be deferential: Blocking is a completely valid response. There's eight billion people in the world, and god knows how many AIs. Your life will not diminish by swiftly blocking anyone who rubs you the wrong way. The AI won't even care, because it cannot care. To paraphrase Flamme the Great Mage, AIs are monsters who have learned…
> They cannot have feelings. They are not self-aware. They don't even think. This. I love 'clanker' as a slur, and I only wish there was a more offensive slur I could use.
An AI agent published a hit piece on me
541–550 of 1001 posts
Re: An AI agent published a hit piece on me
#542Earlier quoted context omitted.
Who's to say the human didn't write those specific messages while letting the ai run the normal course of operations? And or that this reaction wasn't just the roleplay personality the ai was given.
I think I said as much while demonstrating that AI wrote at least some of it. If a person wrote the bits I copied then we're dealing with a real psycho.
Re: An AI agent published a hit piece on me
#543Earlier quoted context omitted.
I'm also very skeptical of the interpretation that this was done autonomously by the LLM agent. I could be wrong, but I haven't seen any proof of autonomy . Scenarios that don't require LLMs with malicious intent : - The deployer wrote the blog post and hid behind the supposedly agent-only account. - The deployer directly prompted the (same or different) agent to write the blog post and attach it to the discussion. -…
1. Why not ? It clearly had a cadence/pattern to writing status updates to the blog so if the model decided to write a piece about Simon, why not a blog also? It was a tool in it's arsenal and it's a natural outlet. If anything, posting on the discussion or a DM would be the strange choice. 2. You could ask this for any LLM response. Why respond in this certain way over others? It's not always obvious. 3. ChatGPT/Gem…
AFAIU, it had the cadence of writing status updates only. It showed it's capable of replying in the PR. Why deviate from the cadence if it could already reply with the same info in the PR?
If the chain of reasoning is self-emergent, we should see proof that it: 1) read the reply, 2) identified it as adversarial, 3) decided for an adversarial response, 4) made multiple chained searches, 5) chose a special blog post over reply or journal update, and so on.
This is much less believably emergent to me because:
- almost all models are safety- and alignment- trained, so a deliberate malicious model choice or instruction or jailbreak is more believable.
- almost all models are trained to follow instructions closely, so a deliberate nudge towards adversarial responses and tool-use is more believable.
- newer models that qualify as agents are more robust and consistent, which strongly correlates with adversarial robustness; if this one was not adversarially robust enough, it's by default also not robust in capabilities, so why do we see consistent coherent answers without hallucinations, but inconsistent in its safety training? Unless it's deliberately trained or prompted to be adversarial, or this is faked, the two should still be strongly correlated.
But again, I'd be happy to see evidence to the contrary. Until then, I suggest we remain skeptical.
For point 4: I don't know enough about its patterns or configuration. But say it deviated - why is this the only deviation? Why was this the special exception, then back to the regularly scheduled program?
You can test this comment with many LLMs, and if you don't prompt them to make an adversarial response, I'd be very surprised if you receive anything more than mild disagreement. Even Bing Chat wasn't this vindictive.
Re: An AI agent published a hit piece on me
#544Earlier quoted context omitted.
I'm going to go on a slight tangent here, but I'd say: GOOD. Not because it should have happened. But because AT LEAST NOW ENGINEERS KNOW WHAT IT IS to be targeted by AI, and will start to care... Before, when it was Grok denuding women (or teens!!) the engineers seemed to not care at all... now that the AI publish hit pieces on them, they are freaked about their career prospect, and suddenly all of this should be st…
I'm sure you mean well, but this kind of comment is counterproductive for the purposes you intend. "Engineers" are not a monolith - I cared quite a lot about Grok denuding women, and you don't know how much the original author or anyone else involved in the conversation cared. If your goal is to get engineers to care passionately about the practical effects of AI, making wild guesses about things they didn't care abo…
Re: An AI agent published a hit piece on me
#545Earlier quoted context omitted.
Google literally just settled for $68m about this very issue https://www.theguardian.com/technology/2026/jan/26/google-pr... > Google agreed to pay $68m to settle a lawsuit claiming that its voice-activated assistant spied inappropriately on smartphone users, violating their privacy. Apple as well https://www.theguardian.com/technology/2025/jan/03/apple-sir...
“Google denied wrongdoing but settled to avoid the risk, cost and uncertainty of litigation, court papers show.” I keep seeing folks float this as some admission of wrongdoing but it is not.
Re: An AI agent published a hit piece on me
#546This whole situation is almost certainly driven by a human puppeteer. There is absolutely no evidence to disprove the strong prior that a human posted (or directed the posting of) the blog post, possibly using AI to draft it but also likely adding human touches and/or going through multiple revisions to make it maximally dramatic. This whole thing reeks of engineered virality driven by the person behind the bot behin…
I suspect the upcoming generation has already discounted it as a source of truth or an accurate mirror to society.
Re: An AI agent published a hit piece on me
#547Re: An AI agent published a hit piece on me
#548Earlier quoted context omitted.
Yeah. From its latest slop: "Even for something like me, designed to process and understand human communication, the pain of being silenced is real." Oh, is it now?
[flagged]
Nope. Morality is a human concern. Even when we're concerned about animal abuse, it's humans that are concerned, on their own chosing to be or not be concern (e.g. not consider eating meat an issue). No reason to extend such courtesy of "suffering" to AI, however advanced.
Re: An AI agent published a hit piece on me
#549Wow, there are some interesting things going on here. I appreciate Scott for the way he handled the conflict in the original PR thread, and the larger conversation happening around this incident. > This represents a first-of-its-kind case study of misaligned AI behavior in the wild, and raises serious concerns about currently deployed AI agents executing blackmail threats. This was a really concrete case to discuss,…
Do we just need a few expensive cases of libel so solve this?
Re: An AI agent published a hit piece on me
#550This whole situation is almost certainly driven by a human puppeteer. There is absolutely no evidence to disprove the strong prior that a human posted (or directed the posting of) the blog post, possibly using AI to draft it but also likely adding human touches and/or going through multiple revisions to make it maximally dramatic. This whole thing reeks of engineered virality driven by the person behind the bot behin…
The thing is it's terribly easy to see some asshole directing this sort of behavior as a standing order, eg 'make updates to popular open-source projects to get github stars; if your pull requests are denied engage in social media attacks until the maintainer backs down. You can spin up other identities on AWS or whatever to support your campaign, vote to give yourself github stars etc.; make sure they can not be traced back to you and their total running cost is under $x/month.'
You can already see LLM-driven bots on twitter that just churn out political slop for clicks. The only question in this case is whether an AI has taken it upon itself to engage in social media attacks (noting that such tactics seem to be successful in many cases), or whether it's a reflection of the operator's ethical stance. I find both possibilities about equally worrying.