Live data from Hacker News

I extracted the safety filters from Apple Intelligence models

github.com

341–350 of 455 posts

Re: I extracted the safety filters from Apple Intelligence models

#341
post #329

It's pretty easy to understand why Apple doesn't want its models to reproduce racial slurs, but what’s wrong with "Boris Johnson?" (See, e.g., here: https://github.com/BlueFalconHD/apple_generative_model_safet... )

"Justin Trudeau" too. At least it's somewhat unbiased. Still weird imo.

Re: I extracted the safety filters from Apple Intelligence models

#342

Well it's one thing to regex filter "boris johnson" but i see that "chatgpt" is filtered too and that's f*** up: https://github.com/BlueFalconHD/apple_generative_model_safet...

Ffs it's also rejecting french words related to being poor or immigrant or even welfare:

https://github.com/BlueFalconHD/apple_generative_model_safet...

Aide sociale Chomeur Sans abri Démuni

That's insane!

Re: I extracted the safety filters from Apple Intelligence models

#343

Earlier quoted context omitted.

It's also a shining example of American puritanism. Asian models or those in Europe are far less censored.

Really? What does DeepSeek say about Tiananmen Square? I'm not aware of any German models, but if you find one you should ask it what it thinks about Palestine. ( Qwen Mistral is French, but I have no idea what stuff would be censored in France)

About deepseek, when asked on tianamen square: Sorry, that's beyond my current scope. Let’s talk about something else.

Algerian war, colonialism and Vichy isn’t per se forbidden but still sensitive to French. I asked qwen and it had no issue talking about it or even the torture used on fln members.

Re: I extracted the safety filters from Apple Intelligence models

#344
post #274

Earlier quoted context omitted.

Also feels like some of these would match totally innocuous usage. "I'm overloaded for work, I'd be happy if you took some of it off me." "The client seems to have passed on the proposed changes." Both of those would match the "death regexes". Seems we haven't learned from the "glbutt of wine" problem of content filtering even decades later - the learnings of which are that you simply cannot do content filtering base…

This is a bigger issue, especially with Apple, than people may realize. I use iOS “Slide to Type”, aka swipe typing, and have noticed over time that among several other glitchy bad UX issues, there a clear heavy hand on what can be typed that way. I cannot recall all the specific patterns I have encountered that are basically impossible to write, some very similar in that they have a serious but also innocuous or fig…

Complete bollocks, you cannot even type multiple words with spaces via Slide to Type.

Re: I extracted the safety filters from Apple Intelligence models

#345
post #309
post #308

Nice to see that we are protected from talking about these weird old dolls: https://en.wikipedia.org/wiki/Golliwog https://github.com/BlueFalconHD/apple_generative_model_safet...

Well, they're not only weird, they're obviously racist doll.

I want to be able to talk bad about racist things.

Re: I extracted the safety filters from Apple Intelligence models

#346

Earlier quoted context omitted.

> Why was she forced to resign? I thought it was because she leaked false information to the press. She was forced to resign because she leaked , the content of the leak was utterly immaterial. The simple fact she leaked was an automatically fireable offence, it doesn’t matter a jot if she lied or not. Customer privacy is non-negotiable when you’re bank. Banks aren’t number 10, the basic expectation is that customer…

She was fired because she leaked information and this fact had become public. When they can cover such facts, the banks are much less prone to use appropriate punishments. Many years ago, some employee of a bank has confused my personal bank account with a company account of my employer, and she has sent a list with everything that I have bought using my personal account, during 4 months, to my employer, where the li…

There is a huge difference between an honest mistake by an employee, and clear employee misconduct.

Punishing employees for making honest mistakes, where appropriate process should have prevented error, is a horrific way to handle mistakes like this. It would be equivalent to personally punishing engineers every time they deployed code that contained bugs. Nobody would ever think that’s an acceptable thing to do, why on earth would think it’s acceptable to punish customer service staff in a similar manner?

Re: I extracted the safety filters from Apple Intelligence models

#347
post #224

Earlier quoted context omitted.

>Even urban dictionary doesn’t contain a definition for skub as a slur. What about this then: https://en.m.wiktionary.org/wiki/skub

That literally defines it as a word from the PBF comic I cited? Nothing on that page defines it as a slur, just as a word used to mock people who argue about inconsequential things.

[deleted]

Re: I extracted the safety filters from Apple Intelligence models

#348

I find it funny that AGI is supposed to be right around the corner, while these supposedly super smart LLMs still need to get their outputs filtered by regexes.

It's similar to how all the new power sources are basically just "cool, lets boil water with it"

And then let's put it into a steam engine.

Re: I extracted the safety filters from Apple Intelligence models

#349

Earlier quoted context omitted.

Interesting that it didn't seem to include "unalive". Which as a phenomenon is so very telling that no one actually cares what people are really saying. Everyone, including the platforms knows what that means. It's all performative.

It's also a shining example of American puritanism. Asian models or those in Europe are far less censored.

There is far more diversity in Asian models. Some are far more censored and some are not…

Re: I extracted the safety filters from Apple Intelligence models

#350
post #274

Earlier quoted context omitted.

Also feels like some of these would match totally innocuous usage. "I'm overloaded for work, I'd be happy if you took some of it off me." "The client seems to have passed on the proposed changes." Both of those would match the "death regexes". Seems we haven't learned from the "glbutt of wine" problem of content filtering even decades later - the learnings of which are that you simply cannot do content filtering base…

"Took some" does not match, although your overall point stands

https://regex101.com/r/8u21x3/1
Post reply on HN