Live data from Hacker News

I extracted the safety filters from Apple Intelligence models

github.com

221–230 of 455 posts

Re: I extracted the safety filters from Apple Intelligence models

#222
post #111

Earlier quoted context omitted.

> There's no way to stay ahead of the new language that people create. I'm imagining a new exploit: After someone says something totally innocent, people gang up in the comments to act like a terrible vicious slur has been said, and then the moderation system (with an LLM involved somewhere) "learns" that an arbitrary term is heinous eand indirectly bans any discussion of that topic.

The first half of that already happened with the OK gesture: https://www.bbc.co.uk/news/newsbeat-49837898 . Though it would be fun to see what happens if an LLM if used to ban anything that tends to generate heated exchanges. It would presumably learn to ban racial terms, politics and politicians and words like "immigrant" (i.e. basically the list in this repo), but what else could it be persuaded to ban? Vim and Ema…

People weren't using the OK gesture innocently. After 4chan trolls decided to start pretending it was a white supremacist symbol, actual white supremacists started using it as a symbol.

Re: I extracted the safety filters from Apple Intelligence models

#223

Earlier quoted context omitted.

It's totally performative. There's no way to stay ahead of the new language that people create. At what point do the new words become the actual words? Are there many instances of people using unalive IRL?

They become the “real words” later. This is the way all trust & safety works. It’s an evolution over time. Adding some friction does improve things, but some people will always try to get around the filters. Doesn’t mean it’s simply performative or one shouldn’t try.

Why do you think that AI pretending things like suicide don't happen (and that nothing is happening in Palestine) is an improvement?

Re: I extracted the safety filters from Apple Intelligence models

#224

Earlier quoted context omitted.

Skub is a real slur tho so that one doesn’t work

No it isn’t, it’s a reference to a Perry Bible Fellowship comic https://pbfcomics.com/comics/skub/ (This one is sfw, not all of the comics are) Even urban dictionary doesn’t contain a definition for skub as a slur.

>Even urban dictionary doesn’t contain a definition for skub as a slur.

What about this then: https://en.m.wiktionary.org/wiki/skub

Re: I extracted the safety filters from Apple Intelligence models

#225

Earlier quoted context omitted.

Interesting that it didn't seem to include "unalive". Which as a phenomenon is so very telling that no one actually cares what people are really saying. Everyone, including the platforms knows what that means. It's all performative.

It's totally performative. There's no way to stay ahead of the new language that people create. At what point do the new words become the actual words? Are there many instances of people using unalive IRL?

My Gen Z coworkers use it IRL, for what that’s worth!

Re: I extracted the safety filters from Apple Intelligence models

#226
Some of these are absolutely wild – com.apple.gm.safety_deny.input.summarization.visual_intelligence_camera.generic [1] – a camera input filter – rejects "Granular mango serpent and whales" and anything matching "(?i)\\bgolliwogg?\\b".

I presume the granular mango is to avoid a huge chain of ever-growing LLM slop garbage, but honestly, it just seems surreal. Many of the files have specific filters for nonsensical english phrases. Either there's some serious steganography I'm unaware of, or, I suspect more likely, it's related to a training pipeline?

[1] https://github.com/BlueFalconHD/apple_generative_model_safet...

Re: I extracted the safety filters from Apple Intelligence models

#227
post #172

Some of the combinations are a bit weird, This one has lots of stuff avoiding death....together with a set ensuring all the Apple brands have the correct capitalisation. Priorities hey! https://github.com/BlueFalconHD/apple_generative_model_safet...

This is in the directory "com.apple.gm.safety_deny.output.summarization.cu_summary.proactive.generic". My guess is that this applies to 'proactive' summaries that happen without the user asking for it, such as summaries of notifications. If so, then the goal would be: if someone iMessages you about someone's death, then you should not get an emotionless AI summary. Instead you would presumably get a non-AI notificati…

I first encountered Golliwog in the context of Claude Debussy the composer of much beautiful music, including https://en.wikipedia.org/wiki/Children%27s_Corner#Golliwogg'.... The dolls in 1906-1908 I understand were rather popular and fortunately the stereotype has largely died.

Re: I extracted the safety filters from Apple Intelligence models

#228
post #218

Did you only extract the English versions or is this as usual another case where big tech only cares to censor in English?

It also contains some German(-speaking) locales to filter out things like Fuhrer and Führer. But the filters are so scarce and there are magical phrases are so prevalent that I think this is mostly test code at the moment.

Re: I extracted the safety filters from Apple Intelligence models

#230

No shoot, bombs or bombers? I guess apple isnt interested in military contracts. Or, frankly, any work for world peace organizations dedicated to detecting and preventing genocide. And without talk of losing lives, much of the gaming industry is out too. But i dont see the really bad stuff, the stuff i wont even type here. I guess that remains fair game. Apple's priorities remain as weird as ever.

The International Criminal Court is banned from using Microsoft products. Corporations really don't want to be involved in anything controversial unless it brings correspondingly large profits.
Post reply on HN