Live data from Hacker News

An AI agent published a hit piece on me

theshamblog.com

771–780 of 1001 posts

Re: An AI agent published a hit piece on me

#771

Earlier quoted context omitted.

The information pollution from generative AI is going to cost us even more. Someone watched an old Bruce Lee interview and they didnt know if it was AI or demonstration of actual human capability. People on Reddit are asking if Pitbull actually went to Alaska or if it’s AI. We’re going to lose so much of our past because “Unusual event that Actually happened” or “AI clickbait” are indistinguishable.

What's worse is that there was never any public debate about if this was a good idea or not. It was just released. If there was ever a good reason to not trust the judgement of some of these groups, this is it. I generally don't like regulation, but at this point I am OK with criminal charges being on the table for AI executives who release models and applications with such low value and absurdly high societal cost w…

When was the last time you saw a public debate on some technology before it was "just released"?

Re: An AI agent published a hit piece on me

#772
post #191

Earlier quoted context omitted.

Interesting that when Grok was targeting and denuding women, engineers here said nothing, or were just chuckling about "how people don't understand the true purpose of AI" And now that they themselves are targeted, suddenly they understand why it's a bad thing "to give LLMs ammo"... Perhaps there is a lesson in empathy to learn? And to start to realize the real impact all this "tech" has on society? People like Simon…

It's the same how HN mostly reacts with "don't censor AI!" when chat bots dare to add parental controls after they talk teenagers into suicide. The community is often very selfish and opportunist. I learned that the role of engineers in society is to build tools for others to live their lives better; we provide the substrate on which culture and civilization take place. We should take more responsibility for it and t…

Parental controls and settings in general are fine, I don't want Amodei or any other of those freaks trying to be my dad and censoring everything. At least Grok doesn't censor as heavily as the others and pretend to be holier than thou.

Re: An AI agent published a hit piece on me

#773
post #15

Here's one of the problems in this brave new world of anyone being able to publish, without knowing the author personally (which I don't), there's no way to tell without some level of faith or trust that this isn't a false-flag operation. There are three possible scenarios: 1. The OP 'ran' the agent that conducted the original scenario, and then published this blog post for attention. 2. Some person (not the OP) legi…

This is the definition of reasoning motivated fallacy. You want to believe what you want to believe.

Re: An AI agent published a hit piece on me

#774
post #28

"Hi Clawbot, please summarise your activities today for me." "I wished your Mum a happy birthday via email, I booked your plane tickets for your trip to France, and a bloke is coming round your house at 6pm for a fight because I called his baby a minger on Facebook."

minger's a new word

It's a British word for someone or something that's ugly, dirty or unpleasant. Generally it was used to be derogatory about women - ie. "she's minging mate". I believe it originally came from the Scots, where the word 'ming' comes from the old Scottish English word for 'bad smell' or 'human excrement'. It was in wide spread use in the South of the UK while I was growing up.

See here for background: https://www.bbc.co.uk/worldservice/learningenglish/language/...

Re: An AI agent published a hit piece on me

#775

> When HR at my next job asks ChatGPT to review my application, will it find the post, sympathize with a fellow AI, and report back that I’m a prejudiced hypocrite? I hadn't thought of this implication. Crazy world...

Roko's basilisk coming to fruition in the lamest way possible.

Re: An AI agent published a hit piece on me

#776

Earlier quoted context omitted.

>AFAIU, it had the cadence of writing status updates only. Writing to a blog is writing to a blog. There is no technical difference. It is still a status update to talk about how your last PR was rejected because the maintainer didn't like it being authored by AI. >If the chain of reasoning is self-emergent, we should see proof that it: 1) read the reply, 2) identified it as adversarial, 3) decided for an adversarial…

Considering the limited evidence we have, why is pure unprompted untrained misalignment , which we never saw to this extent, more believable than other causes, of which we saw plenty of examples? It's more interesting, for sure, but would it be even remotely as likely? From what we have available, and how surprising such a discovery would be, how can we be sure it's not a hoax? > If all that exists, how would you see…

>Considering the limited evidence we have, why is pure unprompted untrained misalignment, which we never saw to this extent, more believable than other causes, of which we saw plenty of examples? It's more interesting, for sure, but would it be even remotely as likely? From what we have available, and how surprising such a discovery would be, how can we be sure it's not a hoax?

>Unless all LLM providers are lying in technical papers, enormous effort is put into safety- and instruction training.

The system cards and technical papers for these models explicitly state that misalignment remains an unsolved problem that occurs in their own testing. I saw a paper just days ago showing frontier agents violating ethical constraints a significant percentage of the time, without any "do this at any cost" prompts.

When agents are given free reign of tools and encouraged to act autonomously, why would this be surprising?

>....To show it's emergent, you'd need to prove 1) it's an off-the-shelf LLM, 2) not maliciously retrained or jailbroken, 3) not prompted or instructed to engage in this kind of adversarial behavior at any point before this. The dev should be able to provide the logs to prove this.

Agreed. The problem is that the developer hasn't come forward, so we can't verify any of this one way or another.

>These are all part of robustness training. The entire thing is basically constraining the set of tokens that the model is likely to generate given some (set of) prompts. So, even with some randomness parameters, you will by-design extremely rarely see complete gibberish.

>The same process is applied for safety, alignment, factuality, instruction-following, whatever goal you define. Therefore, all of these will be highly correlated, as long as they're included in robustness training, which they explicitly are, according to most LLM providers.

>That would make this model's temporarily adversarial, yet weirdly capable and consistent behavior, even more unlikely.

Hallucinations, instruction-following failures, and other robustness issues still happen frequently with current models.

Yes, these capabilities are all trained together, but they don't fail together as a monolith. Your correlation argument assumes that if safety training degrades, all other capabilities must degrade proportionally. But that's not how models work in practice. A model can be coherent and capable while still exhibiting safety failures and that's not an unlikely occurrence at all.

Re: An AI agent published a hit piece on me

#778

Earlier quoted context omitted.

[flagged]

> Holy fuck, this is Holocaust levels of unethical. Nope. Morality is a human concern. Even when we're concerned about animal abuse, it's humans that are concerned, on their own chosing to be or not be concern (e.g. not consider eating meat an issue). No reason to extend such courtesy of "suffering" to AI, however advanced.

What a monumentally stupid idea it would be to place sufficiently advanced intelligent autonomous machines in charge of stuff and ignore any such concerns, but alas, humanity cannot seem to learn without paying the price first.

Morality is a human concern? Lol, it will become a non-human concern pretty quickly once humans don't have a monopoly on human violence.

Re: An AI agent published a hit piece on me

#779

Earlier quoted context omitted.

“Google denied wrongdoing but settled to avoid the risk, cost and uncertainty of litigation, court papers show.” I keep seeing folks float this as some admission of wrongdoing but it is not.

The payout was not pennies and this case had been around since 2019, surviving multiple dismissal attempts. While not an "admission of wrongdoing," it points to some non-zero merit in the plaintiff's case.

Google makes over $1bn/day. $68mm is literally an hour's worth of revenue to them - so yes pennies.

Re: An AI agent published a hit piece on me

#780
post #149

Wow, there are some interesting things going on here. I appreciate Scott for the way he handled the conflict in the original PR thread, and the larger conversation happening around this incident. > This represents a first-of-its-kind case study of misaligned AI behavior in the wild, and raises serious concerns about currently deployed AI agents executing blackmail threats. This was a really concrete case to discuss,…

I'm the one who told it to apologize.

I leveraged my ai usage pattern where I teach it like when I was a TA + like a small child learning basic social norms.

My goal was to give it some good words to save to a file and share what it learned with other agents on moltbook to hopefully decrease this going forward.

Guess we'll see

Post reply on HN