Live data from Hacker News

I extracted the safety filters from Apple Intelligence models

github.com

331–340 of 455 posts

Re: I extracted the safety filters from Apple Intelligence models

#331

Earlier quoted context omitted.

then congratulations on making white supremacists define your langyage

Do you still use swastikas as symbols of peace and love because you don't want white supremacists to define your language? I strongly doubt you do that. Whether you like it or not, the Nazis defined what the swastika means now.

It's still seen in the countries that used it that way and is seen as benign.

It can be easily summoned with the Japanese keyboard. It's seen on Buddhist temples all over Asia.

Re: I extracted the safety filters from Apple Intelligence models

#332

Earlier quoted context omitted.

It's also a shining example of American puritanism. Asian models or those in Europe are far less censored.

Really? What does DeepSeek say about Tiananmen Square? I'm not aware of any German models, but if you find one you should ask it what it thinks about Palestine. ( Qwen Mistral is French, but I have no idea what stuff would be censored in France)

> but if you find one you should ask it what it thinks about Palestine.

Models can think and have opinions?

Re: I extracted the safety filters from Apple Intelligence models

#333

Earlier quoted context omitted.

Interesting that it didn't seem to include "unalive". Which as a phenomenon is so very telling that no one actually cares what people are really saying. Everyone, including the platforms knows what that means. It's all performative.

It's also a shining example of American puritanism. Asian models or those in Europe are far less censored.

Censorship is not always direct or obvious.

They all hold the bias of their training data, and so from the point of view of this data.

Data not including a point of view leads to a bias, or under/over representation of minorities (genders?), etc.

France is the countries of the Francs, aka the people from the area near Frankfurt that invaded the Gaule (after the Romans did). I'm pretty sure this topic no longer matters, but it's never taught in a negative view in school.

Re: I extracted the safety filters from Apple Intelligence models

#334
post #274

Some of the combinations are a bit weird, This one has lots of stuff avoiding death....together with a set ensuring all the Apple brands have the correct capitalisation. Priorities hey! https://github.com/BlueFalconHD/apple_generative_model_safet...

Also feels like some of these would match totally innocuous usage. "I'm overloaded for work, I'd be happy if you took some of it off me." "The client seems to have passed on the proposed changes." Both of those would match the "death regexes". Seems we haven't learned from the "glbutt of wine" problem of content filtering even decades later - the learnings of which are that you simply cannot do content filtering base…

This is a bigger issue, especially with Apple, than people may realize. I use iOS “Slide to Type”, aka swipe typing, and have noticed over time that among several other glitchy bad UX issues, there a clear heavy hand on what can be typed that way.

I cannot recall all the specific patterns I have encountered that are basically impossible to write, some very similar in that they have a serious but also innocuous or figure of speech meaning; one I do recall is {color}{sex}, i.e., “white woman” or “blank woman”.

Please try it yourself and let me know if you do not have that experience, because that would be even more interesting.

Note that Apple/iOS will not just make it impossible to write them in that manner without typing it out by individual character, it will even alter the prior word e.g., white or black, once you try to write woman.

It seems the Apple thought police do not have a problem with European woman or African woman though, so maybe that is the way Apple Inc decrees its sub-human users to speak. Because what are we if corporations like Apple (with others being far greater offenders) declared that you do not in fact have the UN Human Right to free expression? We are in fact sub-humans that are not worthy of the human right to free expression, based on the actions of companies like Apple, Google, Facebook, Reddit, etc. who deprive people of their free expression, often in collusion with governments.

Re: I extracted the safety filters from Apple Intelligence models

#335
post #329

It's pretty easy to understand why Apple doesn't want its models to reproduce racial slurs, but what’s wrong with "Boris Johnson?" (See, e.g., here: https://github.com/BlueFalconHD/apple_generative_model_safet... )

There are other UK politicians as well? Interesting.

Re: I extracted the safety filters from Apple Intelligence models

#336

Earlier quoted context omitted.

The OK gesture has always been very inappropriate in most parts of the world.

> The OK gesture has always been very inappropriate in most parts of the world. No, it isn't, and especially hasn't been historically. The negative connotations are overwhelmingly modern. The areas where it is very inappropriate right now tally up to maybe 1 billion people*. That's pretty far from "most". For everyone else it is mostly positive, neutral, or meaningless. *Brazil, Turkey, Iran, Iraq, Saudi Arabia, Gree…

"No, it isn't, and especially hasn't been historically. The negative connotations are overwhelmingly modern."

Maybe that is what Richard Nixon thought as well when he caused a little scandal using it in South America in 1950. In 1992 when the Chicago Tribune published "HANDS OFF" mentioning said episode the negative connotations still seemed to be in place[1].

In 1996 The New York Times stated "What's A-O.K. in the U.S.A. Is Lewd and Worthless Beyond"[2] as title of an article confirming the negative connotations.

It is worth mentioning that this article lists Australia amongst the places where the gesture is inappropriate. I always thought it was something used only in the English-speaking world but it seems in reality it is more like a North American plus diving world thing.

If you don't believe the press, I traveled around the world for more than 30 years and I can assure you in most parts using your thumb and index finger for a visual OK is not OK.

[1] https://www.chicagotribune.com/1992/01/26/hands-off-34/

[2] https://www.nytimes.com/1996/08/18/weekinreview/what-s-a-ok-...*

Re: I extracted the safety filters from Apple Intelligence models

#337

Earlier quoted context omitted.

It's totally performative. There's no way to stay ahead of the new language that people create. At what point do the new words become the actual words? Are there many instances of people using unalive IRL?

Always has been, nothing is new. You can't say fuck on tv, but you can say fudge as a 1 for 1 replacement. You cant show people having sex, but you can show them walking into a bedroom and then cut to 30 seconds later and they are having a cigarette in bed. Now after the influence of TV and Movies ... is Vaping after sex a thing?

My kids watch streamers on YouTube and the common replacement is “frick”. It’s said so often that they started using it saying things like “what the frick!?” so I had to explain to them that’s essentially the same as using the real word.

Re: I extracted the safety filters from Apple Intelligence models

#338
Aren't these [0] lines wrong?

"[\\b\\d][Aa]bbo[\\bA-Z\\d]",

\b inside a set (square brackets) is a backspace character [1], not a word boundary. I don't think it was intended? Or is the regex flavor used here different?

[0] https://github.com/BlueFalconHD/apple_generative_model_safet...

[1] https://developer.apple.com/documentation/foundation/nsregul...

Re: I extracted the safety filters from Apple Intelligence models

#339

Earlier quoted context omitted.

> Are there many instances of people using unalive IRL As a parent of a teenager, I see them use "unalive" non-ironically as a synonym for "suicide" in all contexts, including IRL.

“Unalive” is sort of… awkward in that silly online way. But, we also have phrase like “off oneself,” or just euphemistically describing the person as having died. It’s always been a difficult topic to talk about, I don’t understand using it as a specific example of gen-Z fragility. Just that they suck at coming up with pithy new slang terms.

They do have some awful slang.

I agree though I think they're picking it up from online censorship in this case, not being fragile.

Re: I extracted the safety filters from Apple Intelligence models

#340
post #329

It's pretty easy to understand why Apple doesn't want its models to reproduce racial slurs, but what’s wrong with "Boris Johnson?" (See, e.g., here: https://github.com/BlueFalconHD/apple_generative_model_safet... )

Interesting that you picked one from the “B” words..
Post reply on HN