Earlier quoted context omitted.
> There's no way to stay ahead of the new language that people create. I'm imagining a new exploit: After someone says something totally innocent, people gang up in the comments to act like a terrible vicious slur has been said, and then the moderation system (with an LLM involved somewhere) "learns" that an arbitrary term is heinous eand indirectly bans any discussion of that topic.
It's not like this unique to LLMs either. By some little trolling on internet you easily can turn hand "OK gesture" into a hate symbol of white supermacy. And fools will fall for it.
I extracted the safety filters from Apple Intelligence models
311–320 of 455 posts
Re: I extracted the safety filters from Apple Intelligence models
#312Earlier quoted context omitted.
> There's no way to stay ahead of the new language that people create. I'm imagining a new exploit: After someone says something totally innocent, people gang up in the comments to act like a terrible vicious slur has been said, and then the moderation system (with an LLM involved somewhere) "learns" that an arbitrary term is heinous eand indirectly bans any discussion of that topic.
It's not like this unique to LLMs either. By some little trolling on internet you easily can turn hand "OK gesture" into a hate symbol of white supermacy. And fools will fall for it.
Re: I extracted the safety filters from Apple Intelligence models
#313Earlier quoted context omitted.
This is the rhetorical tactic of false equivalence. State censorship by an autocracy with the objective of population control is not the same thing as a private company inside a democracy censoring their product to avoid bad press and maintain goodwill for shareholders. If you want solid proof that it's not the same thing, see all the uncensored open weights models that you can freely download and use without fear of…
But who of the general populace has the technical skill to replace their on-device assistant with a free one? And that's if Apple even allows that? In practice, there's not that much difference between a megacorporate monopolist and a state.
Re: I extracted the safety filters from Apple Intelligence models
#314Earlier quoted context omitted.
Apple's 1984 ad is so hypocritical today. This is Apple actively steering public thought. No code - anywhere - should look like this. I don't care if the politicians are right, left, or authoritarian. This is wrong.
Why is this wrong? Applying special treatment to politically exposed persons has been standard practice in every high risk industry for a very long time. The simple fact is that people get extremely emotional about politicians, politicians both receive obscene amounts of abuse, and have repeatedly demonstrated they’re not above weaponising tools like this for their own goals. Seems perfectly reasonable that Apple doe…
Not that getting the latest trash talk is the main vocation of pretrained AIs anyway.
The only risk here is that some third grade journalist of a third grade newspaper writes another article about how outrageous some generated AI statement is. An article that should be completely ignored instead of it leading to more censorship.
And Apple flinches here, so in the end it means it cannot provide a sensible general model. It would be affected by their censorship.
Re: I extracted the safety filters from Apple Intelligence models
#315Earlier quoted context omitted.
Interesting that it didn't seem to include "unalive". Which as a phenomenon is so very telling that no one actually cares what people are really saying. Everyone, including the platforms knows what that means. It's all performative.
It's also a shining example of American puritanism. Asian models or those in Europe are far less censored.
Re: I extracted the safety filters from Apple Intelligence models
#316Earlier quoted context omitted.
This question is sort of the same as asking why the universal translator wasn't able to translate the metaphor language of the Star Trek episode Darmok. Surely if the metaphor has become the first order meaning then there's no litteral meaning anymore.
The only reason kids started using "unalive" is to get around Youtube filters that disallow the use of the word "kill"
Re: I extracted the safety filters from Apple Intelligence models
#317Some of the combinations are a bit weird, This one has lots of stuff avoiding death....together with a set ensuring all the Apple brands have the correct capitalisation. Priorities hey! https://github.com/BlueFalconHD/apple_generative_model_safet...
Also feels like some of these would match totally innocuous usage. "I'm overloaded for work, I'd be happy if you took some of it off me." "The client seems to have passed on the proposed changes." Both of those would match the "death regexes". Seems we haven't learned from the "glbutt of wine" problem of content filtering even decades later - the learnings of which are that you simply cannot do content filtering base…
Re: I extracted the safety filters from Apple Intelligence models
#318Earlier quoted context omitted.
People weren't using the OK gesture innocently. After 4chan trolls decided to start pretending it was a white supremacist symbol, actual white supremacists started using it as a symbol.
then congratulations on making white supremacists define your langyage
I strongly doubt you do that. Whether you like it or not, the Nazis defined what the swastika means now.
Re: I extracted the safety filters from Apple Intelligence models
#319Earlier quoted context omitted.
Also feels like some of these would match totally innocuous usage. "I'm overloaded for work, I'd be happy if you took some of it off me." "The client seems to have passed on the proposed changes." Both of those would match the "death regexes". Seems we haven't learned from the "glbutt of wine" problem of content filtering even decades later - the learnings of which are that you simply cannot do content filtering base…
"Took some" does not match, although your overall point stands