Live data from Hacker News

I extracted the safety filters from Apple Intelligence models

github.com

171–180 of 455 posts

Re: I extracted the safety filters from Apple Intelligence models

#171

Earlier quoted context omitted.

Hey I was pro-skub waaaay before all the anti-skub people switched sides.

Skub is a real slur tho so that one doesn’t work

No it isn’t, it’s a reference to a Perry Bible Fellowship comic https://pbfcomics.com/comics/skub/

(This one is sfw, not all of the comics are)

Even urban dictionary doesn’t contain a definition for skub as a slur.

Re: I extracted the safety filters from Apple Intelligence models

#172

Some of the combinations are a bit weird, This one has lots of stuff avoiding death....together with a set ensuring all the Apple brands have the correct capitalisation. Priorities hey! https://github.com/BlueFalconHD/apple_generative_model_safet...

This is in the directory "com.apple.gm.safety_deny.output.summarization.cu_summary.proactive.generic".

My guess is that this applies to 'proactive' summaries that happen without the user asking for it, such as summaries of notifications.

If so, then the goal would be: if someone iMessages you about someone's death, then you should not get an emotionless AI summary. Instead you would presumably get a non-AI notification showing the full text or a truncated version of the text.

In other words, avoid situations like this story [1], where someone found it "dystopian" to get an Apple Intelligence summary of messages in which someone broke up with them.

For that use case, filtering for death seems entirely appropriate, though underinclusive.

This filter doesn’t seem to apply when you explicitly request a summary of some text using Writing Tools. That probably corresponds to “com.apple.gm.safety_deny.output.summarization.text_assistant.generic” [2], which has a different filter that only rejects two things: "Granular mango serpent", and "golliwogg".

Sure enough, I was able to get Writing Tools to give me summaries containing "death", but in cases where the summary should contain "granular mango serpent" or "golliwogg", I instead get an error saying "Writing Tools aren't designed to work with this type of content." (Actually that might be the input filter rather than the output filter; whatever.)

"Granular mango serpent" is probably a test case that's meant to be unlikely to appear in real documents. Compare to "xylophone copious opportunity defined elephant" from the code_intelligence safety filter, where the first letter of each word spells out "Xcode".

But one might ask what's so special about "golliwogg". It apparently refers to an old racial caricature, but why is that the one and only thing that needs filtering?

[1] https://arstechnica.com/ai/2024/10/man-learns-hes-being-dump...

[2] https://github.com/BlueFalconHD/apple_generative_model_safet...

Re: I extracted the safety filters from Apple Intelligence models

#173

Earlier quoted context omitted.

Interesting that it didn't seem to include "unalive". Which as a phenomenon is so very telling that no one actually cares what people are really saying. Everyone, including the platforms knows what that means. It's all performative.

It's totally performative. There's no way to stay ahead of the new language that people create. At what point do the new words become the actual words? Are there many instances of people using unalive IRL?

> Are there many instances of people using unalive IRL?

In my experience yes. This is already commonplace. Mostly, but not exclusively, amongst the younger generation.

Re: I extracted the safety filters from Apple Intelligence models

#174

Earlier quoted context omitted.

Well that’s sad. They can’t even face the word ?

It’s not about whether they can face it. The younger generations are more in tune with mental health and topics like suicide than any previous generation. The etymology of the euphemism was about avoiding online censorship, while its “IRL” usage was merely absorbed through familiarity from the online usage.

But unalive self is suicide and unalive is just death, right? For example, You can unalive other people against their will...

Re: I extracted the safety filters from Apple Intelligence models

#175
post #162

Earlier quoted context omitted.

Unalive and other self censors were adopted by young people because the tiktok algorithm would reprioritize videos that included specific words. Then it made its way into the culture. It has nothing to do with being performative

I think what they meant is that the platforms are being performative by attempting to crack down on those specific words. If saying "killed" is not allowed but "unalived" is permitted and the users all agree that they mean the same thing, then the ban on the word "killed" doesn't accomplish anything.

What does using the grape emoji when talking about sexual assault accomplish? I see videos, compassionate, kind people who make videos speaking to victims in a completely serious tone use this emoji.

People talk about tiktok algorithm on tiktok. I don't even know...

Re: I extracted the safety filters from Apple Intelligence models

#176

Earlier quoted context omitted.

I'm not surprised that anything political is being filtered, but this should definitely provoke some deep consideration around who has control of this stuff.

You’re not wrong, and it’s something we “doomers” have been saying since OpenAI dumped ChatGPT onto folks. These are curated walled gardens, and everyone should absolutely be asking what ulterior motives are in play for the owners of said products.

Some of us really value offline and uncensored LLMs for this and more reasons, but that doesn’t solve the problem it just reduces or changes the bias.

Re: I extracted the safety filters from Apple Intelligence models

#177

Earlier quoted context omitted.

I can Google for any of these people, and I can get real results with real information.

You would hope that search would be a politically safe space to operate. But politicians find a way to ruin everything for short term political gain. https://arstechnica.com/tech-policy/2018/12/republicans-in-c...

I would hope!

But no one actually believes Google is politically neutral do they?

Re: I extracted the safety filters from Apple Intelligence models

#178
post #51

Earlier quoted context omitted.

I’m telling you, some people have weird fantasies…

Now that they've cleaned it up it isn't so bad, but browse Civit.ai a bit and that'll still be confirmed - just not with real people anymore.

I’m convinced there are a dozen deviants on Covid with a hundred new accounts per month posting their perversion in order to make it seem more commonplace.

No porn site has that much extremely X or Y stuff.

Someone is using the internets newest porn site to push a sexual agenda.

Re: I extracted the safety filters from Apple Intelligence models

#179

Some of the combinations are a bit weird, This one has lots of stuff avoiding death....together with a set ensuring all the Apple brands have the correct capitalisation. Priorities hey! https://github.com/BlueFalconHD/apple_generative_model_safet...

Interesting that it didn't seem to include "unalive". Which as a phenomenon is so very telling that no one actually cares what people are really saying. Everyone, including the platforms knows what that means. It's all performative.

Good, let them. Don't give them a reason to crack down on speech.
Post reply on HN