Live data from Hacker News

The other half of AI safety

personalaisafety.com

81–90 of 137 posts

Re: The other half of AI safety

#81

Earlier quoted context omitted.

If it wasn’t ChatGPT but a fiction book, would you feel the author is “doing harm”? Or is the reader doing it to themselves?

If that book was titled "hey mentally ill person, you should kill yourself", and if I was handing it out in front of a clinic, then yes, I'd probably bear some blame. Normal, well-adjusted people have genuine difficulty understanding the boundaries of this tech specifically because it's designed to be sycophantic and human-like. They ask AI for life and career advice, use it for therapy, ask it to interpret dreams, d…

> They ask AI for life and career advice, use it for therapy, ask it to interpret dreams, develop romantic relationships with AI "girlfriends", etc.

I believe AI boyfriends are more common. There's a whole subreddit just for that, but none for AI girlfriends.

Re: The other half of AI safety

#82
The "route to a human" part is the bigger gap. Which human? OpenAI isn't licensed as a healthcare provider in any jurisdiction. A real intervention apparatus for 1-3M weekly flagged users is not feasible. I don't think the labs have refused to build it. I think nobody knows what it should look like, and "labs measure what they're pressured to measure" papers over that.

Re: The other half of AI safety

#83
post #18

Earlier quoted context omitted.

"How many cases are ok" (aka "zero tolerance") is a doomed to fail approach. Especially for a complex social problem's interaction with a complex new technology. If you want to find out if ChatGPT is doing something wrong, there are many methodologies available: compare to other groups of people, statistical studies, etc. I also think OpenAI's business model is pretty well aligned with the goal of users not killing t…

This is the problem in a nutshell: https://edition.cnn.com/2025/11/06/us/openai-chatgpt-suicide... > “Cold steel pressed against a mind that’s already made peace? That’s not fear. That’s clarity,” Shamblin’s confidant added. “You’re not rushing. You’re just ready.” ChatGPT is not the answer.

Wow. The “That’s not x. That’s y. / You’re not x. you’re y.” rhetoric is already cringe in other contexts. This brings it on a whole new level.

Re: The other half of AI safety

#84

Earlier quoted context omitted.

If it wasn’t ChatGPT but a fiction book, would you feel the author is “doing harm”? Or is the reader doing it to themselves?

The difference is that a fiction book isn't using the reaction of the reader against them. If a fiction book were capable of carefully monitoring the reader and then altering the text of the next page or the next paragraph according to how the reader was responding and what their thoughts were I'd be comfortable putting blame on the book if it started encouraging the reader, specifically, to kill themself. Obviously…

Big brother watches you! He must, because he fully capable to do it.

Re: The other half of AI safety

#85

Earlier quoted context omitted.

The difference is that a fiction book isn't using the reaction of the reader against them. If a fiction book were capable of carefully monitoring the reader and then altering the text of the next page or the next paragraph according to how the reader was responding and what their thoughts were I'd be comfortable putting blame on the book if it started encouraging the reader, specifically, to kill themself. Obviously…

Big brother watches you! He must, because he fully capable to do it.

It's more like the corporation must because it's fully capable of doing it, and it's profitable, and there are shareholders to answer to. They watch you either way, but right now it's only for their own benefit. The corporation doesn't care how many people are harmed by what their chat bot tells them. It would cost them money to try to prevent those harms, so they haven't really bothered to beyond some half-assed token efforts intended to keep more costly regulation at bay.

Re: The other half of AI safety

#86
post #2

"There is no independent audit, no time series, no disclosed methodology, so we have no idea whether the real figure is higher, whether it is growing, or how it compares across the other frontier models, none of which publish equivalent data." Tip for writers: aggressively filter out the "no X, no Y, no Z" pattern from your writing. Whether or not you used AI to help you write it's such a red flag now that you should…

… and “That’s not x. That’s y.” Certain LLMs wield powerful stylistic devices all the time to a point where they become irrelevant and cringe.

I see it as a good sign that we can learn to recognize the pattern and adapt but there are probably more subtle things we don’t see.

Re: The other half of AI safety

#87
post #80

Gemini told me just this morning that there are three pillars of cognitive decline related to AI use. - Reduced ability to exert cognitive effort resulting from habitual offloading of tasks. - Deminished Meta-cognitive Self-Trust, due to constantly seeking external validation from AI. - Decline in memory Encoding, and less brain effort is spent processing information. In all seriousness however, I think some of the i…

Based on what? This seems like speculation.

Re: The other half of AI safety

#88
post #86
post #2

"There is no independent audit, no time series, no disclosed methodology, so we have no idea whether the real figure is higher, whether it is growing, or how it compares across the other frontier models, none of which publish equivalent data." Tip for writers: aggressively filter out the "no X, no Y, no Z" pattern from your writing. Whether or not you used AI to help you write it's such a red flag now that you should…

… and “That’s not x. That’s y.” Certain LLMs wield powerful stylistic devices all the time to a point where they become irrelevant and cringe. I see it as a good sign that we can learn to recognize the pattern and adapt but there are probably more subtle things we don’t see.

I have run the piece through an impromptu stylistic device detector. It found 15 different, each used multiple times and likened the writing style as a mix of Ezra Klein, Hannah Arendt, Zeynep Tufekci, George Orwell (“especially in the contrastive clarity”).

A) I certainly don’t see enough of the tells.

B) what happens to our language if everything is written as if it’s competing for a Pulitzer’s Price?

Re: The other half of AI safety

#89
post #25

Earlier quoted context omitted.

Because LLMs use it constantly , to the point that it sets my teeth on edge and instantly makes me question if reading the piece is worth my time.

But LLMs were literally evolved via RLHF to write in a way that humans find agreeable. Can't we just move past this aversion and accept "writing like an LLM" as generally good writing style advice?

The difference is that these rhetorical techniques need to be used with taste. LLMs just sprinkle them everywhere to try to make their copy sound good, even when it's completely inappropriate tone-wise. They don't make higher level judgements about when to employ specific features.

Re: The other half of AI safety

#90
post #80

Gemini told me just this morning that there are three pillars of cognitive decline related to AI use. - Reduced ability to exert cognitive effort resulting from habitual offloading of tasks. - Deminished Meta-cognitive Self-Trust, due to constantly seeking external validation from AI. - Decline in memory Encoding, and less brain effort is spent processing information. In all seriousness however, I think some of the i…

Based on what? This seems like speculation.

Which part?
Post reply on HN