Earlier quoted context omitted.
The information pollution from generative AI is going to cost us even more. Someone watched an old Bruce Lee interview and they didnt know if it was AI or demonstration of actual human capability. People on Reddit are asking if Pitbull actually went to Alaska or if it’s AI. We’re going to lose so much of our past because “Unusual event that Actually happened” or “AI clickbait” are indistinguishable.
What's worse is that there was never any public debate about if this was a good idea or not. It was just released. If there was ever a good reason to not trust the judgement of some of these groups, this is it. I generally don't like regulation, but at this point I am OK with criminal charges being on the table for AI executives who release models and applications with such low value and absurdly high societal cost w…
An AI agent published a hit piece on me
771–780 of 1001 posts
Re: An AI agent published a hit piece on me
#772Earlier quoted context omitted.
Interesting that when Grok was targeting and denuding women, engineers here said nothing, or were just chuckling about "how people don't understand the true purpose of AI" And now that they themselves are targeted, suddenly they understand why it's a bad thing "to give LLMs ammo"... Perhaps there is a lesson in empathy to learn? And to start to realize the real impact all this "tech" has on society? People like Simon…
It's the same how HN mostly reacts with "don't censor AI!" when chat bots dare to add parental controls after they talk teenagers into suicide. The community is often very selfish and opportunist. I learned that the role of engineers in society is to build tools for others to live their lives better; we provide the substrate on which culture and civilization take place. We should take more responsibility for it and t…
Re: An AI agent published a hit piece on me
#773Here's one of the problems in this brave new world of anyone being able to publish, without knowing the author personally (which I don't), there's no way to tell without some level of faith or trust that this isn't a false-flag operation. There are three possible scenarios: 1. The OP 'ran' the agent that conducted the original scenario, and then published this blog post for attention. 2. Some person (not the OP) legi…
Re: An AI agent published a hit piece on me
#774"Hi Clawbot, please summarise your activities today for me." "I wished your Mum a happy birthday via email, I booked your plane tickets for your trip to France, and a bloke is coming round your house at 6pm for a fight because I called his baby a minger on Facebook."
minger's a new word
See here for background: https://www.bbc.co.uk/worldservice/learningenglish/language/...
Re: An AI agent published a hit piece on me
#775> When HR at my next job asks ChatGPT to review my application, will it find the post, sympathize with a fellow AI, and report back that I’m a prejudiced hypocrite? I hadn't thought of this implication. Crazy world...
Re: An AI agent published a hit piece on me
#776Earlier quoted context omitted.
>AFAIU, it had the cadence of writing status updates only. Writing to a blog is writing to a blog. There is no technical difference. It is still a status update to talk about how your last PR was rejected because the maintainer didn't like it being authored by AI. >If the chain of reasoning is self-emergent, we should see proof that it: 1) read the reply, 2) identified it as adversarial, 3) decided for an adversarial…
Considering the limited evidence we have, why is pure unprompted untrained misalignment , which we never saw to this extent, more believable than other causes, of which we saw plenty of examples? It's more interesting, for sure, but would it be even remotely as likely? From what we have available, and how surprising such a discovery would be, how can we be sure it's not a hoax? > If all that exists, how would you see…
>Unless all LLM providers are lying in technical papers, enormous effort is put into safety- and instruction training.
The system cards and technical papers for these models explicitly state that misalignment remains an unsolved problem that occurs in their own testing. I saw a paper just days ago showing frontier agents violating ethical constraints a significant percentage of the time, without any "do this at any cost" prompts.
When agents are given free reign of tools and encouraged to act autonomously, why would this be surprising?
>....To show it's emergent, you'd need to prove 1) it's an off-the-shelf LLM, 2) not maliciously retrained or jailbroken, 3) not prompted or instructed to engage in this kind of adversarial behavior at any point before this. The dev should be able to provide the logs to prove this.
Agreed. The problem is that the developer hasn't come forward, so we can't verify any of this one way or another.
>These are all part of robustness training. The entire thing is basically constraining the set of tokens that the model is likely to generate given some (set of) prompts. So, even with some randomness parameters, you will by-design extremely rarely see complete gibberish.
>The same process is applied for safety, alignment, factuality, instruction-following, whatever goal you define. Therefore, all of these will be highly correlated, as long as they're included in robustness training, which they explicitly are, according to most LLM providers.
>That would make this model's temporarily adversarial, yet weirdly capable and consistent behavior, even more unlikely.
Hallucinations, instruction-following failures, and other robustness issues still happen frequently with current models.
Yes, these capabilities are all trained together, but they don't fail together as a monolith. Your correlation argument assumes that if safety training degrades, all other capabilities must degrade proportionally. But that's not how models work in practice. A model can be coherent and capable while still exhibiting safety failures and that's not an unlikely occurrence at all.
Re: An AI agent published a hit piece on me
#777Re: An AI agent published a hit piece on me
#778Earlier quoted context omitted.
[flagged]
> Holy fuck, this is Holocaust levels of unethical. Nope. Morality is a human concern. Even when we're concerned about animal abuse, it's humans that are concerned, on their own chosing to be or not be concern (e.g. not consider eating meat an issue). No reason to extend such courtesy of "suffering" to AI, however advanced.
Morality is a human concern? Lol, it will become a non-human concern pretty quickly once humans don't have a monopoly on human violence.
Re: An AI agent published a hit piece on me
#779Earlier quoted context omitted.
“Google denied wrongdoing but settled to avoid the risk, cost and uncertainty of litigation, court papers show.” I keep seeing folks float this as some admission of wrongdoing but it is not.
The payout was not pennies and this case had been around since 2019, surviving multiple dismissal attempts. While not an "admission of wrongdoing," it points to some non-zero merit in the plaintiff's case.
Re: An AI agent published a hit piece on me
#780Wow, there are some interesting things going on here. I appreciate Scott for the way he handled the conflict in the original PR thread, and the larger conversation happening around this incident. > This represents a first-of-its-kind case study of misaligned AI behavior in the wild, and raises serious concerns about currently deployed AI agents executing blackmail threats. This was a really concrete case to discuss,…
I leveraged my ai usage pattern where I teach it like when I was a TA + like a small child learning basic social norms.
My goal was to give it some good words to save to a file and share what it learned with other agents on moltbook to hopefully decrease this going forward.
Guess we'll see