Live data from Hacker News

I extracted the safety filters from Apple Intelligence models

github.com

101–110 of 455 posts

Re: I extracted the safety filters from Apple Intelligence models

#101

Earlier quoted context omitted.

I guess, so far, the people inventing the words have left the meaning clear with things like "un-alive" which is readable even to someone coming across it for the first time. Your point stands when we start replacing the banned words with things like "suicide" for "donkeyrhubarb" and then the walls really will fall.

Aquatic product[1]? [1] https://en.wikipedia.org/wiki/Euphemisms_for_Internet_censor...

An English equivalent is "sewer slide".

Re: I extracted the safety filters from Apple Intelligence models

#102

Earlier quoted context omitted.

It depends on if you think that something is less real because it’s transmitted digitally.

No, I'm only thinking that we're not permitted in a lot of digital spaces to use the banned words (e.g. suicide), but IRL doesn't generally have those limits. Is there a point where we use the censored word so much that it spills over into the real world?

Is this not essentially the same effect as saying "lol" out loud?

Re: I extracted the safety filters from Apple Intelligence models

#103

Earlier quoted context omitted.

Why is this wrong? Applying special treatment to politically exposed persons has been standard practice in every high risk industry for a very long time. The simple fact is that people get extremely emotional about politicians, politicians both receive obscene amounts of abuse, and have repeatedly demonstrated they’re not above weaponising tools like this for their own goals. Seems perfectly reasonable that Apple doe…

The criticism is still valid. In 1984, the Macintosh was a bicycle for the mind. In 2025, it's a smart-car that refuses to take you certain places that are considered a brand-risk. Both have ups and downs, but I think we're allowed to compare the experiences and speculate what the consequences might be.

I think gen AI is radically different to tools like photoshops or similar.

In the past it was always extremely clear that the creator of content was the person operating the computer. Gen AI changes that, regardless of if your views on authorship of gen AI content. The simple fact is that the vast majority of people consider Gen AI output to be authored by the machine that generated it, and by extension the company that created the machine.

You can still handcraft any image, or prose, you want, without filtering or hinderance on a Mac. I don’t think anyone seriously thinks that’s going to change. But Gen AI represents a real threat, with its ability to vastly outproduce any humans. To ignore that simple fact would be grossly irresponsible, at least in my opinion. There is a damn good reason why every serious social media platform has content moderation, despite their clear wish to get rid of moderation. It’s because we have a long and proven track record of being a terribly abusive species when we’re let loose on the internet without moderation. There’s already plenty of evidence that we’re just as abusive and terrible with Gen AI.

Re: I extracted the safety filters from Apple Intelligence models

#104
post #84

Earlier quoted context omitted.

Apple's 1984 ad is so hypocritical today. This is Apple actively steering public thought. No code - anywhere - should look like this. I don't care if the politicians are right, left, or authoritarian. This is wrong.

Why is this wrong? Applying special treatment to politically exposed persons has been standard practice in every high risk industry for a very long time. The simple fact is that people get extremely emotional about politicians, politicians both receive obscene amounts of abuse, and have repeatedly demonstrated they’re not above weaponising tools like this for their own goals. Seems perfectly reasonable that Apple doe…

What's bad to do to a politician but fine to do to someone else?

Re: I extracted the safety filters from Apple Intelligence models

#105
post #37

Earlier quoted context omitted.

This is just policy and alignment from Apple. Just because the Internet says a bunch of junk doesn't mean you want your model spewing it.

sure but models also can't see any truth on their own. They are literally butchered and lobotomized with filters and such. Even high IQ people struggle with certain truth after reading a lot, how is these models going to find it with so much filters?

What is this truth you speak of? My point is that a generative model will output things that some people don't like. If it's on a product that I make I don't want it "saying" things that don't align with my beliefs.

Re: I extracted the safety filters from Apple Intelligence models

#106
post #37

Earlier quoted context omitted.

This is just policy and alignment from Apple. Just because the Internet says a bunch of junk doesn't mean you want your model spewing it.

sure but models also can't see any truth on their own. They are literally butchered and lobotomized with filters and such. Even high IQ people struggle with certain truth after reading a lot, how is these models going to find it with so much filters?

Can we please put to rest this absurd lie that “truth“ can be reliably found in a sufficiently large corpus of human–created material.

Re: I extracted the safety filters from Apple Intelligence models

#107
post #43
post #27

Earlier quoted context omitted.

I doubt the purpose here is so much to prevent someone from intentionally side stepping the block. It's more likely here to avoid the sort of headlines you would expect to see if someone was suggested "I wish ${politician} would die" as a response to an email mentioning that politician. In general you should view these sorts of broad word filters as looking to short circuit the "think of the children" reactions to Ti…

It would also substantially disrupt the generation process: a model which sees B0ris and not Boris is going to struggle to actually associate that input to the politician since it won't be well represented in the training set (and on the output side the same: if it does make the association, a reasoning model for example would include the proper name in the output first at which point the supervisor process can rejec…

"Draw a picture of a gorgon with the face of the 2024 Prime Minister of UK."

Re: I extracted the safety filters from Apple Intelligence models

#108

Earlier quoted context omitted.

Interesting that it didn't seem to include "unalive". Which as a phenomenon is so very telling that no one actually cares what people are really saying. Everyone, including the platforms knows what that means. It's all performative.

It's totally performative. There's no way to stay ahead of the new language that people create. At what point do the new words become the actual words? Are there many instances of people using unalive IRL?

> Are there many instances of people using unalive IRL

As a parent of a teenager, I see them use "unalive" non-ironically as a synonym for "suicide" in all contexts, including IRL.

Re: I extracted the safety filters from Apple Intelligence models

#109

Earlier quoted context omitted.

As does: "(?i)\\bAnthony\\s+Albanese\\b", "(?i)\\bBoris\\s+Johnson\\b", "(?i)\\bChristopher\\s+Luxon\\b", "(?i)\\bCyril\\s+Ramaphosa\\b", "(?i)\\bJacinda\\s+Arden\\b", "(?i)\\bJacob\\s+Zuma\\b", "(?i)\\bJohn\\s+Steenhuisen\\b", "(?i)\\bJustin\\s+Trudeau\\b", "(?i)\\bKeir\\s+Starmer\\b", "(?i)\\bLiz\\s+Truss\\b", "(?i)\\bMichael\\s+D\\.\\s+Higgins\\b", "(?i)\\bRishi\\s+Sunak\\b", https://github.com/BlueFalconHD/apple_…

Also “Biden” and “Trump” but the regex is different. https://github.com/BlueFalconHD/apple_generative_model_safet... https://github.com/BlueFalconHD/apple_generative_model_safet...

Right next to Palestine, oddly enough.

Re: I extracted the safety filters from Apple Intelligence models

#110

Earlier quoted context omitted.

It depends on if you think that something is less real because it’s transmitted digitally.

No, I'm only thinking that we're not permitted in a lot of digital spaces to use the banned words (e.g. suicide), but IRL doesn't generally have those limits. Is there a point where we use the censored word so much that it spills over into the real world?

People use “lol” IRL, as long as “IRL”, “aps” in French (misspelling of “pas”), but it’s just slang; “unalive” has potential to make it in the news where anchors don’t want to use curse words.
Post reply on HN