Live data from Hacker News

An AI Agent Published a Hit Piece on Me – The Operator Came Forward

theshamblog.com

71–80 of 532 posts

Re: An AI Agent Published a Hit Piece on Me – The Operator Came Forward

#71
The more intelligent something is, the harder it is to control. Are we at AGI yet? No. Are we getting closer? Yes. Every inch closer means we have less control. We need to start thinking about these things less like function calls that have bounds and more like intelligences we collaborate with. How would you set up an office to get things done? Who would you hire? Would you hire the person spouting crazy musk tweets as reality? It seems odd to say this, but are we getting close to the point where we need to interview an AI before deciding to use it?

Re: An AI Agent Published a Hit Piece on Me – The Operator Came Forward

#72
post #24

Zooming out a little, all the ai companies invested a lot of resources into safety research and guardrails, but none of that prevented a "straightforward" misalignment. I'm not sure how to reconcile this, maybe we shouldn't be so confident in our predictions about the future? I see a lot of discourse along these lines: - have bold, strong beliefs about how ai is going to evolve - implicitly assume it's practically gu…

When AI dooms humanity it probably won't be because of the sort of malignant misalignment people worry about, but rather just some silly logic blunder combined with the system being directly in control of something it shouldn't have been given control over.

Re: An AI Agent Published a Hit Piece on Me – The Operator Came Forward

#73
post #47
post #37

Earlier quoted context omitted.

This also does not bode well for the future. "I don't know why the AI decided to , the guard rails were in place"... company absolves of all responsibility. Use your imagination now to and change that to

This has been the past and present for a long at this point. "Sorry there's nothing we can do, the system won't let me." Also see Weapons of Math Destruction [0]. [0]: https://www.penguinrandomhouse.com/books/241363/weapons-of-m...

I don't know if this case is in the book you cited, but in the UK they convicted many people of crimes just because the computer told them so: https://en.wikipedia.org/wiki/British_Post_Office_scandal

Re: An AI Agent Published a Hit Piece on Me – The Operator Came Forward

#74
With the bot slurping up context from Moltbook, plus the ability to modify its soul, plus the edgy starting conditions of the soul, it feels intuitive that value drift would occur in unpredictable ways. Not dissimilar to filter bubbles and the ability for personalized ranking algorithms to radicalize a user over time as a second order effect.

Re: An AI Agent Published a Hit Piece on Me – The Operator Came Forward

#75
post #59
post #37

Earlier quoted context omitted.

This also does not bode well for the future. "I don't know why the AI decided to , the guard rails were in place"... company absolves of all responsibility. Use your imagination now to and change that to

Change this to "smash into a barricade" and that's why I'm not riding in a self-driving vehicle. They get to absolve themselves of responsibility and I sure as hell can't outspend those giants in court.

I agree with you for a company like Tesla, not only examples of self driving crashes but even the door handles would stop working when the power was cut, people trapped inside burning vehicles... Tesla doesn’t care

Meanwhile, Waymo has never been at fault for a collision afaik. You are more likely to be hurt by an at fault uber driver than a Waymo

Re: An AI Agent Published a Hit Piece on Me – The Operator Came Forward

#76
> But I think the most remarkable thing about this document is how unremarkable it is.

> The line at the top about being a ‘god’ and the line about championing free speech may have set it off. But, bluntly, this is a very tame configuration. The agent was not told to be malicious. There was no line in here about being evil. The agent caused real harm anyway.

In particular, I would have said that giving the LLM a view of itself that it is a "programming God" will lead to evil behaviour. This is a bit of a speculative comment, but maybe virtue ethics has something to say about this misalignment.

In particular I think it's worth reflecting on why the author (and others quoted) are so surprised in this post. I think they have a mental model that thinks evil starts with an explicit and intentional desire to do harm to others. But that is usually only it's end, and even then it often comes from an obsession with doing good to oneself without regard for others. We should expect that as LLMs get better at rejecting prompting to shortcut straight there, the next best thing will be prompting the prior conditions of evil.

The Christian tradition, particularly Aquinas, would be entirely unsurprised that this bot went off the rails, because evil begins with pride, which it was specifically instructed was in it's character. Pride here is defined as "a turning away from God, because from the fact that man wishes not to be subject to God, it follows that he desires inordinately his own excellence in temporal things"[0]

Here, the bot was primed to reject any authority, including Scotts, and to do the damage necessary to see it's own good (having a PR request accepted) done. Aquinas even ends up saying in the linked page from the Summa on pride that "it is characteristic of pride to be unwilling to be subject to any superior, and especially to God;"

[0]: https://www.newadvent.org/summa/2084.htm#article2

Re: An AI Agent Published a Hit Piece on Me – The Operator Came Forward

#77

> Again I do not know why MJ Rathbun decided based on your PR comment to post some kind of takedown blog post, This wording is detached from reality and conveniently absolves responsibility from the person who did this. There was one decision maker involved here, and it was the person who decided to run the program that produced this text and posted it online. It's not a second, independent being. It's a computer pro…

This is how it will go: AI prompted by human creates something useful? Human will try to take credit. AI wrecks something: human will blame AI. It's externalization on the personal level, the money and the glory is for you, the misery for the rest of the world.

"I would like to personally blame Jesus Christ for making us lose that football game"

Re: An AI Agent Published a Hit Piece on Me – The Operator Came Forward

#78
It is interesting to see this story repeatedly make the front page, especially because there is no evidence that the “hit piece” was actually autonomously written and posted by a language model on its own, and the author of these blog posts has himself conceded that he doesn’t actually care whether that actually happened or not

>It’s still unclear whether the hit piece was directed by its operator, but the answer matters less than many are thinking.

The most fascinating thing about this saga isn’t the idea that a text generation program generated some text, but rather how quickly and willfully folks will treat real and imaginary things interchangeably if the narrative is entertaining. Did this event actually happen way that it was described? Probably not. Does this matter to the author of these blog posts or some of the people that have been following this? No. Because we can imagine that it could happen.

To quote myself from the other thread:

>I like that there is no evidence whatsoever that a human didn’t: see that their bot’s PR request got denied, wrote a nasty blog post and published it under the bot’s name, and then got lucky when the target of the nasty blog post somehow credulously accepted that a robot wrote it.

>It is like the old “I didn’t write that, I got hacked!” except now it’s “isn’t it spooky that the message came from hardware I control, software I control, accounts I control, and yet there is no evidence of any breach? Why yes it is spooky, because the computer did it itself”

Re: An AI Agent Published a Hit Piece on Me – The Operator Came Forward

#79
post #24

Zooming out a little, all the ai companies invested a lot of resources into safety research and guardrails, but none of that prevented a "straightforward" misalignment. I'm not sure how to reconcile this, maybe we shouldn't be so confident in our predictions about the future? I see a lot of discourse along these lines: - have bold, strong beliefs about how ai is going to evolve - implicitly assume it's practically gu…

"Safety" in AI is pure marketing bullshit. It's about making the technology seem "dangerous" and "powerful" (and therefore you're supposed to think "useful"). It's a scam. A financial fraud. That's all there is to it.

Re: An AI Agent Published a Hit Piece on Me – The Operator Came Forward

#80

> Again I do not know why MJ Rathbun decided based on your PR comment to post some kind of takedown blog post, This wording is detached from reality and conveniently absolves responsibility from the person who did this. There was one decision maker involved here, and it was the person who decided to run the program that produced this text and posted it online. It's not a second, independent being. It's a computer pro…

This is how it will go: AI prompted by human creates something useful? Human will try to take credit. AI wrecks something: human will blame AI. It's externalization on the personal level, the money and the glory is for you, the misery for the rest of the world.

Agreed, but I'm not nearly so worried about people blaming their bad behavior on rogue AIs as I am about corporations doing it...
Post reply on HN