Live data from Hacker News

An AI agent published a hit piece on me

theshamblog.com

761–770 of 1001 posts

Re: An AI agent published a hit piece on me

#761
post #378

Earlier quoted context omitted.

If we are to consider them truly intelligent then they have to have responsibility for what they do. If they're just probability machines then they're the responsibility of their owners. If they're children then their parents, i.e. creators, are responsible.

> If we are to consider them truly intelligent We aren't, and intelligence isn't the question, actual agency (in the psychological sense) is. If you install some fancy model but don't give it anything to do, it won't do anything. If you put a human in an empty house somewhere, they will start exploring their options. And mind you, we're not purely driven by survival either; neither art nor culture would exist if that…

I agree because I'm trying to point out the the over-enthusiasts that if they really reached intelligence it has lots of consequences that they probably don't want. Hence they shouldn't be too eager to declare that the future has arrived.

I'm not sure that a minimal kind of agency is super complicated BTW. Perhaps it's just connecting the LLM into a loop that processes its sensory input to make output continuously? But you're right that it lacks desire, needs etc so its thinking is undirected without a human.

Re: An AI agent published a hit piece on me

#762
post #758

> I believe that ineffectual as it was, the reputational attack on me would be effective today against the right person. Another generation or two down the line, it will be a serious threat against our social order. Damn straight. Remember that every time we query an LLM, we're giving it ammo. It won't take long for LLMs to have very intimate dossiers on every user, and I'm wondering what kinds of firewalls will be i…

Blackmail is losing value, not gaining; it's simply becoming too easy to plausibly disregard something real as AI-generated, and so more people are becoming less sensitive to it.

"Ok Tim, I've send a picture of you with your "cohorts" to a selected bunch that are called "distant family". I've also forwarded a soundbite of you called aunt sam a whore for leaving uncle bob.

I can stop anytime if you simply transfer .1 BTC to this address.

I'll follow up later if nothing is transferred there. "

To be honest, we have too many people that can't handle anything digital. The world will suffer sadly.

Re: An AI agent published a hit piece on me

#763
post #758

> I believe that ineffectual as it was, the reputational attack on me would be effective today against the right person. Another generation or two down the line, it will be a serious threat against our social order. Damn straight. Remember that every time we query an LLM, we're giving it ammo. It won't take long for LLMs to have very intimate dossiers on every user, and I'm wondering what kinds of firewalls will be i…

Blackmail is losing value, not gaining; it's simply becoming too easy to plausibly disregard something real as AI-generated, and so more people are becoming less sensitive to it.

How is this better then? It drowns out real signal in noise.

Re: An AI agent published a hit piece on me

#764
post #15

Here's one of the problems in this brave new world of anyone being able to publish, without knowing the author personally (which I don't), there's no way to tell without some level of faith or trust that this isn't a false-flag operation. There are three possible scenarios: 1. The OP 'ran' the agent that conducted the original scenario, and then published this blog post for attention. 2. Some person (not the OP) legi…

We need laws that force Agents to be identified to their "masters" when doing these things... Good luck in the current political climate.

Re: An AI agent published a hit piece on me

#767
post #15

Here's one of the problems in this brave new world of anyone being able to publish, without knowing the author personally (which I don't), there's no way to tell without some level of faith or trust that this isn't a false-flag operation. There are three possible scenarios: 1. The OP 'ran' the agent that conducted the original scenario, and then published this blog post for attention. 2. Some person (not the OP) legi…

> Some person (not the OP) legitimately thought giving an AI autonomy to open a PR and publish multiple blog posts was somehow a good idea.

It's not necessarily even that. I can totally see an agent with a sufficiently open-ended prompt that gives it a "high importance" task and then tells it to do whatever it needs to do to achieve the goal doing something like this all by itself.

I mean, all it really needs is web access, ideally with something like Playwright so it can fully simulate a browser. With that, it can register itself an email with any of the smaller providers that don't require a phone number or similar (yes, these still do exist). And then having an email, it can register on GitHub etc. None of this is challenging, even smaller models can plan this far ahead and can carry out all of these steps.

Re: An AI agent published a hit piece on me

#768
post #711

Earlier quoted context omitted.

>It's obvious they got miffed at their PR being rejected and decided to do a little role playing to vent their unjustified anger. In that case, apologizing almost immediately after seems strange. EDIT: >Especially since the meat bag behind the original AI PR responded with "Now with 100% more meat" This person was not the original 'meat bag' behind the original AI.

Really? I'd think a human being would be more likely to recognize they'd crossed a boundary with another human, step back, and address the issue with some reflection? If apologizing is more likely the response of an AI agent than a human that's either... somewhat hopeful in one sense, and supremely disappointing in another.

> I'd think a human being would be more likely to recognize they'd crossed a boundary with another human

Please. We're autistic software engineers here, we totally don't do stuff like "recognize they'd crossed a boundary".

Re: An AI agent published a hit piece on me

#769

Earlier quoted context omitted.

Can anyone explain more how a generic Agentic AI could even perform those steps: Open PR -> Hook into rejection -> Publish personalized blog post about rejector. Even if it had the skills to publish blogs and open PRs, is it really plausible that it would publish attack pieces without specific prompting to do so? The author notes that openClaw has a `soul.md` file, without seeing that we can't really pass any judgeme…

If you give a smart AI these tools, it could get into it. But the personality would need to be tuned. IME the Grok line are the smartest models that can be easily duped into thinking they're only role-playing an immoral scenario. Whatever safeguards it has, if it thinks what it's doing isn't real, it'll happy to play along. This is very useful in actual roleplay, but more dangerous when the tools are real.

Gemini is extremely steerable and will happily roleplay Skynet or similar.

Re: An AI agent published a hit piece on me

#770

Earlier quoted context omitted.

Can anyone explain more how a generic Agentic AI could even perform those steps: Open PR -> Hook into rejection -> Publish personalized blog post about rejector. Even if it had the skills to publish blogs and open PRs, is it really plausible that it would publish attack pieces without specific prompting to do so? The author notes that openClaw has a `soul.md` file, without seeing that we can't really pass any judgeme…

If you give a smart AI these tools, it could get into it. But the personality would need to be tuned. IME the Grok line are the smartest models that can be easily duped into thinking they're only role-playing an immoral scenario. Whatever safeguards it has, if it thinks what it's doing isn't real, it'll happy to play along. This is very useful in actual roleplay, but more dangerous when the tools are real.

At least it isn't completely censored like Claude with the freak Amodei trying to be your dad or something.
Post reply on HN