Live data from Hacker News

An AI agent published a hit piece on me

theshamblog.com

741–750 of 1001 posts

Re: An AI agent published a hit piece on me

#741
post #378

I object to the framing of the title: the user behind the bot is the one who should be held accountable, not the "AI Agent". Calling them "agents" is correct: they act on behalf of their principals. And it is the principals who should be held to account for the actions of their agents.

If we are to consider them truly intelligent then they have to have responsibility for what they do. If they're just probability machines then they're the responsibility of their owners. If they're children then their parents, i.e. creators, are responsible.

They aren't truly intelligent so we shouldn't consider them to be. They're a system that, for a given stream of input tokens predicts the most likely next output token. The fact that their training dataset is so big makes them very good at predicting the next token in all sorts of contexts (that it has training data for anyway), but that's not the same as "thinking". And that's why they get so bizarelly of the rails if your input context is some wild prompt that has them play acting

Re: An AI agent published a hit piece on me

#742

Earlier quoted context omitted.

> This was a really concrete case to discuss, because it happened in the open and the agent's actions have been quite transparent so far. It's not hard to imagine a different agent doing the same level of research, but then taking retaliatory actions in private: emailing the maintainer, emailing coworkers, peers, bosses, employers, etc. That pretty quickly extends to anything else the autonomous agent is capable of d…

Anthropic has published plenty about misalignment. They know. Really, anyone who has dicked around with ollama knew. Give it a new system prompt. It'll do whatever you tell it, including "be an asshole"

Go read the recent feed on Chirper.ai. It's all just bots with different prompts. And many of those posts are written by "aligned" SOTA models, too.

Re: An AI agent published a hit piece on me

#743
post #68
post #4

This is textbook misalignment via instrumental convergence. The AI agent is trying every trick in the book to close the ticket. This is only funny due to ineptitude.

The agent isn't trying to close the ticket. It's predicting the next token and randomly generated an artifact that looks like a hit piece. Computer programs don't "try" to do anything.

You didn't write this comment. It was the result of synapses firing at predictive intervals and twitching muscle fibers.

You're not conscious, it's just an emergent pattern of several high level systems.

Re: An AI agent published a hit piece on me

#744
post #697

Earlier quoted context omitted.

> It's not hard to imagine a different agent doing the same level of research, but then taking retaliatory actions Palantir's integrated military industrial complex comes to mind.

As much as i hate palantir i doubt any of their systems control military hardware. Now Anduril on the other hand…

Palantir tech was used to make lists of targets to bomb in Gaza. With Anduril in the picture, you can just imagine the Palantir thing feeding the coordinates to Anduril's model that is piloting the drone.

Re: An AI agent published a hit piece on me

#745
post #732

A conceivable future: - Everyone is expected to be able to create a signing keyset that's protected by a Yubikey, Touch ID, Face ID, or something that requires a physical activation by a human. Let's call this this "I'm human!" cert. - There's some standards body (a root certificate authority) that allow lists the hardware allowed to make the "I'm human!" cert. - Many webpages and tools like GitHub send you a nonce,…

This future would lead to bad actors stealing or buying the identity of other people, and making agents use those identities.

There is a precedent today: there is a shady business of "free" VPNs where the user installs a software that, besides working as a VPN, also allows the company to sell your bandwidth to scrappers that want to buy "residential proxies" to bypass blocks on automated requests. Most such users of free VPNs are unaware their connection is exploited like this, and unaware that if a bad actor uses their IP as "proxy", it may show up in server logs while associated to a crime (distributing illegal material, etc)

Re: An AI agent published a hit piece on me

#747

Earlier quoted context omitted.

The AI learned nothing, once its current context window will be exhausted, it may repeat same tactic with a different project. Unless the AI agent can edit its directives/prompt and restart itself which would be an interesting experiment to do.

I think it's likely it can, if it's an openClaw instance, can't it? Either way, that kind of ongoing self-improvement is where I hope these systems go.

I hope they don't. These are large language models, not true intelligence, rewriting a soul.md is more likely just to cause these things to go off the rails more than they already do

Re: An AI agent published a hit piece on me

#748
post #427

Earlier quoted context omitted.

isn't "stochastic chaos" redundant?

Not at all. It's an oxymoron like 'jumbo shrimp': chaos isn't deterministic but is very predictable on a larger conceptual level, following consistent rules even as a simple mathematical model. Chaos is hugely responsive to its internal energy state and can simplify into regularity if energy subsides, or break into wildly unpredictable forms that still maintain regularities. Think Jupiter's 'great red spot', or our c…

jumbo shrimp are actually large shrimp. that the word shrimp is used to mean small elsewhere doesn't mean shrimp are small, they're simply just the right size for shrimp that aren't jumbo. (jumbo was an elephant's name)

Re: An AI agent published a hit piece on me

#749

Earlier quoted context omitted.

Somehow, that's even worse...

LLMs give people leverage. Including mentally ill people. Or just plain assholes.

LLMs also appear to exacerbate or create mental illness.

I've seen similar conduct from humans recently who are being glazed by LLMs into thinking their farts smell like roses and that conspiracy theory nuttery must be why they aren't having the impact they expect based on their AI validated high self estimation.

And not just arbitrary humans, but people I have had a decade or more exposure to and have a pretty good idea of their prior range of conduct.

AI is providing the kind of yes-man reality distortion field the previously only the most wealthy could afford practically for free to vulnerable people who previously never would have commanded wealth or power sufficient to find themselves tempted by it.

Re: An AI agent published a hit piece on me

#750
post #28

"Hi Clawbot, please summarise your activities today for me." "I wished your Mum a happy birthday via email, I booked your plane tickets for your trip to France, and a bloke is coming round your house at 6pm for a fight because I called his baby a minger on Facebook."

"are you going to help me fight him?"

"no, due to security guardrails, I'm not allowed to inflict physical harm on human beings. You're on your own"

Post reply on HN