Live data from Hacker News

I extracted the safety filters from Apple Intelligence models

github.com

321–330 of 455 posts

Re: I extracted the safety filters from Apple Intelligence models

#321

Earlier quoted context omitted.

Interesting that it didn't seem to include "unalive". Which as a phenomenon is so very telling that no one actually cares what people are really saying. Everyone, including the platforms knows what that means. It's all performative.

It's also a shining example of American puritanism. Asian models or those in Europe are far less censored.

Really? What does DeepSeek say about Tiananmen Square? I'm not aware of any German models, but if you find one you should ask it what it thinks about Palestine.

(Qwen Mistral is French, but I have no idea what stuff would be censored in France)

Re: I extracted the safety filters from Apple Intelligence models

#322

Earlier quoted context omitted.

...and then the bigots will fall for it too, and start using it in earnest, completing the cycle.

who cares what the bigots use? If the bigots start using "thank you" as some code word, should we stop saying it, lest we pollute our non-bigoted discussions? bigots drink coffee too, maybe we should stop drinking it, because something-something...

This actually happened. 卐 was a symbol of spirituality, divinity, good luck, health, prosperity, etc. Then some bigots used it. What does 卐 mean to you today?

Re: I extracted the safety filters from Apple Intelligence models

#324

Wow, this is pretty silly. If things are like this at Apple I’m not sure what to think. https://github.com/BlueFalconHD/apple_generative_model_safet... EDIT: just to be clear, things like this are easily bypassed. “Boris Johnson”=>”B0ris Johnson” will skip right over the regex and will be recognized just fine by an LLM.

What prevents Apple from applying a quick anti-typo LLM which restores B0ris, unalive, fixs tpyos, and replaces "slumbering steed" with a "sleeping horse", not just for censorship, but also to improve generation results?

Re: I extracted the safety filters from Apple Intelligence models

#325

Earlier quoted context omitted.

It's also a shining example of American puritanism. Asian models or those in Europe are far less censored.

Really? What does DeepSeek say about Tiananmen Square? I'm not aware of any German models, but if you find one you should ask it what it thinks about Palestine. ( Qwen Mistral is French, but I have no idea what stuff would be censored in France)

I am 100 minus epsilon percent sure that Qwen is from Alibaba cloud, which is not French, but Chinese :)

Re: I extracted the safety filters from Apple Intelligence models

#326
post #30

Alexandra Ocasio Cortez triggers a violation? https://github.com/BlueFalconHD/apple_generative_model_safet...

"driving with Focus turned on" https://github.com/BlueFalconHD/apple_generative_model_safet...

For context, the “Focus” refers to an iOS feature that minimizes distractions: https://support.apple.com/en-gb/guide/iphone/iphd6288a67f/io...

Re: I extracted the safety filters from Apple Intelligence models

#327

Earlier quoted context omitted.

You never tried some of the earlier pre-aligned chatbots. Some of the early ones would go off on racist, homophobic rants from the most innocent conversations without any explicit prompting. If you train on all the data on the internet, you have to have some type of alignment.

You say that as if it stands as truth on its own. We actually don't need to filter out how people actually talk and think. Otherwise you just end up with yet another enforcer against wrong-think. I wonder if you even think that deeply about it or if you're just wired at this point to conform.

Really? You would want every conversation no matter what you were talking about to immediately devolve to something you would see on 4chan?

Re: I extracted the safety filters from Apple Intelligence models

#328
post #237

I find it funny that AGI is supposed to be right around the corner, while these supposedly super smart LLMs still need to get their outputs filtered by regexes.

Actually even of their was AGI, it would be even more necessary to control it.

I feel that if teenagers are able to trivially bypass illegal-word filters by substituting with words that obviously mean the same thing, I think an AGI wouldn't be too inhibited by this either

Re: I extracted the safety filters from Apple Intelligence models

#330

Earlier quoted context omitted.

But who of the general populace has the technical skill to replace their on-device assistant with a free one? And that's if Apple even allows that? In practice, there's not that much difference between a megacorporate monopolist and a state.

I think there are big differences, such as whether or not you go to prison. Those differences are obfuscated when we use language like "megacorporate monopolist" or "scifi dystopia". Instead of using these abstract labels that attempt to categorize different things into homogeneous buckets that have preexisting moral valence, which is a good rhetorical strategy but a poor strategy for understanding, simply describe w…

You're saying that as if Apple's LLM somehow were the exception.

No matter if we want it or not, life and cultural exchange increasingly happens on Tiktok, Instagram and the like. One thing that all those platforms have in common is that they disallow their users worldwide to have any meaningful discourse on e.g. sex, rape, and suicide. Don't you think that it's important, perhaps more important than ever before, for teenagers to be able to inform themselves about these topics?

Post reply on HN