Live data from Hacker News

I extracted the safety filters from Apple Intelligence models

github.com

311–320 of 455 posts

Re: I extracted the safety filters from Apple Intelligence models

#311
post #198
post #111

Earlier quoted context omitted.

> There's no way to stay ahead of the new language that people create. I'm imagining a new exploit: After someone says something totally innocent, people gang up in the comments to act like a terrible vicious slur has been said, and then the moderation system (with an LLM involved somewhere) "learns" that an arbitrary term is heinous eand indirectly bans any discussion of that topic.

It's not like this unique to LLMs either. By some little trolling on internet you easily can turn hand "OK gesture" into a hate symbol of white supermacy. And fools will fall for it.

[deleted]

Re: I extracted the safety filters from Apple Intelligence models

#312
post #198
post #111

Earlier quoted context omitted.

> There's no way to stay ahead of the new language that people create. I'm imagining a new exploit: After someone says something totally innocent, people gang up in the comments to act like a terrible vicious slur has been said, and then the moderation system (with an LLM involved somewhere) "learns" that an arbitrary term is heinous eand indirectly bans any discussion of that topic.

It's not like this unique to LLMs either. By some little trolling on internet you easily can turn hand "OK gesture" into a hate symbol of white supermacy. And fools will fall for it.

That reminds me of a question I have since I saw my first LLM hallucination: How much do people think hallucination/confabulation can be attributed to trolling and sarcasm having slipped into the training data? Is it possible we could get the rate of hallucinations down by better filtering of cynicism from the traing data?

Re: I extracted the safety filters from Apple Intelligence models

#313

Earlier quoted context omitted.

This is the rhetorical tactic of false equivalence. State censorship by an autocracy with the objective of population control is not the same thing as a private company inside a democracy censoring their product to avoid bad press and maintain goodwill for shareholders. If you want solid proof that it's not the same thing, see all the uncensored open weights models that you can freely download and use without fear of…

But who of the general populace has the technical skill to replace their on-device assistant with a free one? And that's if Apple even allows that? In practice, there's not that much difference between a megacorporate monopolist and a state.

So in modern times, not being able to generate an image of suicide on your phone whenever you want means you are suffering from communist censorship?

Re: I extracted the safety filters from Apple Intelligence models

#314
post #84

Earlier quoted context omitted.

Apple's 1984 ad is so hypocritical today. This is Apple actively steering public thought. No code - anywhere - should look like this. I don't care if the politicians are right, left, or authoritarian. This is wrong.

Why is this wrong? Applying special treatment to politically exposed persons has been standard practice in every high risk industry for a very long time. The simple fact is that people get extremely emotional about politicians, politicians both receive obscene amounts of abuse, and have repeatedly demonstrated they’re not above weaponising tools like this for their own goals. Seems perfectly reasonable that Apple doe…

What do you mean reasonable? I know that some Apple users tend to outsource "possibilities" to their favorite company, but I would obviously want an AI to not be affected by the political bitching du jours.

Not that getting the latest trash talk is the main vocation of pretrained AIs anyway.

The only risk here is that some third grade journalist of a third grade newspaper writes another article about how outrageous some generated AI statement is. An article that should be completely ignored instead of it leading to more censorship.

And Apple flinches here, so in the end it means it cannot provide a sensible general model. It would be affected by their censorship.

Re: I extracted the safety filters from Apple Intelligence models

#315

Earlier quoted context omitted.

Interesting that it didn't seem to include "unalive". Which as a phenomenon is so very telling that no one actually cares what people are really saying. Everyone, including the platforms knows what that means. It's all performative.

It's also a shining example of American puritanism. Asian models or those in Europe are far less censored.

[flagged]

Re: I extracted the safety filters from Apple Intelligence models

#316

Earlier quoted context omitted.

This question is sort of the same as asking why the universal translator wasn't able to translate the metaphor language of the Star Trek episode Darmok. Surely if the metaphor has become the first order meaning then there's no litteral meaning anymore.

The only reason kids started using "unalive" is to get around Youtube filters that disallow the use of the word "kill"

Pretty sure TikTok filters do the same and was also a major influence in using that term

Re: I extracted the safety filters from Apple Intelligence models

#317
post #274

Some of the combinations are a bit weird, This one has lots of stuff avoiding death....together with a set ensuring all the Apple brands have the correct capitalisation. Priorities hey! https://github.com/BlueFalconHD/apple_generative_model_safet...

Also feels like some of these would match totally innocuous usage. "I'm overloaded for work, I'd be happy if you took some of it off me." "The client seems to have passed on the proposed changes." Both of those would match the "death regexes". Seems we haven't learned from the "glbutt of wine" problem of content filtering even decades later - the learnings of which are that you simply cannot do content filtering base…

"Took some" does not match, although your overall point stands

Re: I extracted the safety filters from Apple Intelligence models

#318

Earlier quoted context omitted.

People weren't using the OK gesture innocently. After 4chan trolls decided to start pretending it was a white supremacist symbol, actual white supremacists started using it as a symbol.

then congratulations on making white supremacists define your langyage

Do you still use swastikas as symbols of peace and love because you don't want white supremacists to define your language?

I strongly doubt you do that. Whether you like it or not, the Nazis defined what the swastika means now.

Re: I extracted the safety filters from Apple Intelligence models

#319
post #274

Earlier quoted context omitted.

Also feels like some of these would match totally innocuous usage. "I'm overloaded for work, I'd be happy if you took some of it off me." "The client seems to have passed on the proposed changes." Both of those would match the "death regexes". Seems we haven't learned from the "glbutt of wine" problem of content filtering even decades later - the learnings of which are that you simply cannot do content filtering base…

"Took some" does not match, although your overall point stands

"off me"

Re: I extracted the safety filters from Apple Intelligence models

#320

Earlier quoted context omitted.

It's also a shining example of American puritanism. Asian models or those in Europe are far less censored.

I'm sure this has more to do with legal liability than morals.

Which is a reflection of morality, of sorts.
Post reply on HN