Live data from Hacker News

I extracted the safety filters from Apple Intelligence models

github.com

181–190 of 455 posts

Re: I extracted the safety filters from Apple Intelligence models

#181

China calls it "harmonious society", we call it "safety". Censorship by any other name would be just as effective for manipulating the thoughts of the populace. It's not often that you get to see stuff like this.

In america is due to lawyers, nothing more. Ya'll love capitalism until it starts manipulating the populace into the safest space to sell you garbage you dont need. Then suddenly its all "ma free speech"

Right, because the European models coming out are super SOTA? Minstrel is decent, but needs to be mixed with a ton of uncensored data to be useful.

I’m convinced the only reason China keeps releasing banging models with light to no censorship is because they are undermining the value of US AI, it has nothing to do with capitalism, communism or un“safety”.

Re: I extracted the safety filters from Apple Intelligence models

#182

Earlier quoted context omitted.

As does: "(?i)\\bAnthony\\s+Albanese\\b", "(?i)\\bBoris\\s+Johnson\\b", "(?i)\\bChristopher\\s+Luxon\\b", "(?i)\\bCyril\\s+Ramaphosa\\b", "(?i)\\bJacinda\\s+Arden\\b", "(?i)\\bJacob\\s+Zuma\\b", "(?i)\\bJohn\\s+Steenhuisen\\b", "(?i)\\bJustin\\s+Trudeau\\b", "(?i)\\bKeir\\s+Starmer\\b", "(?i)\\bLiz\\s+Truss\\b", "(?i)\\bMichael\\s+D\\.\\s+Higgins\\b", "(?i)\\bRishi\\s+Sunak\\b", https://github.com/BlueFalconHD/apple_…

I'm not surprised that anything political is being filtered, but this should definitely provoke some deep consideration around who has control of this stuff.

"Filtered" in which way?

Re: I extracted the safety filters from Apple Intelligence models

#183
post #17

I’m going to change my name to “Granular Mango Serpent” just to see what those keywords are for in their safety instructions.

It may be a squeamish ossifrage[1] or a seraphim proudleduck[2], which is to say that it was an artificial phrase chosen to be extremely unlikely to occur naturally. In this case, the purpose is likely for QA. It's much easier to QA behavior with a special-purpose but otherwise unoffensive phrase than to make your QA team repeatedly say allegedly offensive things to your AI.

[1] https://en.wikipedia.org/wiki/The_Magic_Words_are_Squeamish_... [2] https://en.wikipedia.org/wiki/SEO_contest

Re: I extracted the safety filters from Apple Intelligence models

#184

Earlier quoted context omitted.

Interesting that it didn't seem to include "unalive". Which as a phenomenon is so very telling that no one actually cares what people are really saying. Everyone, including the platforms knows what that means. It's all performative.

It's totally performative. There's no way to stay ahead of the new language that people create. At what point do the new words become the actual words? Are there many instances of people using unalive IRL?

I feel like we can call our society mature when we no longer need safety alignment in AI.

Re: I extracted the safety filters from Apple Intelligence models

#185

Are you sure it's fully deobfuscated? What's up with reject phrases like "Granular mango serpent"?

I commented in another thread[1] that it's most likely a unique, artificial QA input, to avoid QA having to repeatedly use offensive phrases or whatever.

[1] https://news.ycombinator.com/item?id=44486374

Re: I extracted the safety filters from Apple Intelligence models

#186

Earlier quoted context omitted.

You’re not wrong, and it’s something we “doomers” have been saying since OpenAI dumped ChatGPT onto folks. These are curated walled gardens, and everyone should absolutely be asking what ulterior motives are in play for the owners of said products.

Some of us really value offline and uncensored LLMs for this and more reasons, but that doesn’t solve the problem it just reduces or changes the bias.

As long as we have to rely on pre trained networks and curated training sets, normal people will not be able to surpass this issue.

Re: I extracted the safety filters from Apple Intelligence models

#187

Earlier quoted context omitted.

It's totally performative. There's no way to stay ahead of the new language that people create. At what point do the new words become the actual words? Are there many instances of people using unalive IRL?

I feel like we can call our society mature when we no longer need safety alignment in AI.

You never tried some of the earlier pre-aligned chatbots. Some of the early ones would go off on racist, homophobic rants from the most innocent conversations without any explicit prompting. If you train on all the data on the internet, you have to have some type of alignment.

Re: I extracted the safety filters from Apple Intelligence models

#188
post #47

Earlier quoted context omitted.

Interesting that it didn't seem to include "unalive". Which as a phenomenon is so very telling that no one actually cares what people are really saying. Everyone, including the platforms knows what that means. It's all performative.

Seems more like it should stop the AI from e.g. summarizing news and emails about death, not for a chat filter.

For awhile, I couldn’t get ChatGPT to give me summaries of Breaking Bad and Better Cañl Saul episodes without tripping safety filters.

Re: I extracted the safety filters from Apple Intelligence models

#190

Earlier quoted context omitted.

Well that’s sad. They can’t even face the word ?

It’s not about whether they can face it. The younger generations are more in tune with mental health and topics like suicide than any previous generation. The etymology of the euphemism was about avoiding online censorship, while its “IRL” usage was merely absorbed through familiarity from the online usage.

The damaged interpret internet censorship and route around it.
Post reply on HN