Live data from Hacker News

I extracted the safety filters from Apple Intelligence models

github.com

191–200 of 455 posts

Re: I extracted the safety filters from Apple Intelligence models

#191

Earlier quoted context omitted.

I feel like we can call our society mature when we no longer need safety alignment in AI.

You never tried some of the earlier pre-aligned chatbots. Some of the early ones would go off on racist, homophobic rants from the most innocent conversations without any explicit prompting. If you train on all the data on the internet, you have to have some type of alignment.

You say that as if it stands as truth on its own. We actually don't need to filter out how people actually talk and think. Otherwise you just end up with yet another enforcer against wrong-think. I wonder if you even think that deeply about it or if you're just wired at this point to conform.

Re: I extracted the safety filters from Apple Intelligence models

#192
Here's a combined file of all the non-locale-specific rules, for easier review: https://github.com/BlueFalconHD/apple_generative_model_safet...

It was generated as part of this PR to consolidate the metadata.json files: https://github.com/BlueFalconHD/apple_generative_model_safet...

Re: I extracted the safety filters from Apple Intelligence models

#193

Earlier quoted context omitted.

I don't think it's controversial or unsurprising at all that a company doesn't want their random sentence generator to spit out 'brand damaging' sentences. You know the field day media would have Apple's new feature summarises a text message as "Jane thinks Anthony Albanese should die".

When the choice is between 1. "avoid tarnishing my own brand" and 2. "doing what the user requested," corporations will always choose option 1. Who is this software supposed to be serving, anyway? I'm surprised MS Office still allows me to type "Microsoft can go suck a dick" into a document and Apple's Pages app still allows me to type "Apple are hypocritical jerks." I wonder how long until that won't be the case...

> I wonder how long until that won't be the case...

when there's no more alternative word processors any more.

Re: I extracted the safety filters from Apple Intelligence models

#194
post #92

Earlier quoted context omitted.

As does: "(?i)\\bAnthony\\s+Albanese\\b", "(?i)\\bBoris\\s+Johnson\\b", "(?i)\\bChristopher\\s+Luxon\\b", "(?i)\\bCyril\\s+Ramaphosa\\b", "(?i)\\bJacinda\\s+Arden\\b", "(?i)\\bJacob\\s+Zuma\\b", "(?i)\\bJohn\\s+Steenhuisen\\b", "(?i)\\bJustin\\s+Trudeau\\b", "(?i)\\bKeir\\s+Starmer\\b", "(?i)\\bLiz\\s+Truss\\b", "(?i)\\bMichael\\s+D\\.\\s+Higgins\\b", "(?i)\\bRishi\\s+Sunak\\b", https://github.com/BlueFalconHD/apple_…

They spelled Jacinda Ardern's name wrong.

Just in the region/CN file, weirdly.

Re: I extracted the safety filters from Apple Intelligence models

#195

Earlier quoted context omitted.

> Are there many instances of people using unalive IRL As a parent of a teenager, I see them use "unalive" non-ironically as a synonym for "suicide" in all contexts, including IRL.

Well that’s sad. They can’t even face the word ?

I think it's just the term they immediately associate with the idea. They see "unalive" more than "suicide" online, so it becomes their default word for it. The fact that it originates in automated censorship avoidance is irrelevant.

Re: I extracted the safety filters from Apple Intelligence models

#196
post #17

I’m going to change my name to “Granular Mango Serpent” just to see what those keywords are for in their safety instructions.

It may be a squeamish ossifrage[1] or a seraphim proudleduck[2], which is to say that it was an artificial phrase chosen to be extremely unlikely to occur naturally. In this case, the purpose is likely for QA. It's much easier to QA behavior with a special-purpose but otherwise unoffensive phrase than to make your QA team repeatedly say allegedly offensive things to your AI. [1] https://en.wikipedia.org/wiki/The_Magi…

I think the EICAR test file [1] is more apt. Rather than passing around actually malicious files as part of your tests, it's better to just have it recognize an innocuous and unlikely pattern as malware.

[1] https://en.wikipedia.org/wiki/EICAR_test_file

Re: I extracted the safety filters from Apple Intelligence models

#197

Earlier quoted context omitted.

Interesting that it didn't seem to include "unalive". Which as a phenomenon is so very telling that no one actually cares what people are really saying. Everyone, including the platforms knows what that means. It's all performative.

yo, these are businesses. It's not performative, its CYA. They care because of legal reasons, not moral or ethical.

yo, so it's a performance they're putting on as a legal fig leaf, rather than a genuine attempt to prevent people talking about the concept of death?

Re: I extracted the safety filters from Apple Intelligence models

#198
post #111

Earlier quoted context omitted.

It's totally performative. There's no way to stay ahead of the new language that people create. At what point do the new words become the actual words? Are there many instances of people using unalive IRL?

> There's no way to stay ahead of the new language that people create. I'm imagining a new exploit: After someone says something totally innocent, people gang up in the comments to act like a terrible vicious slur has been said, and then the moderation system (with an LLM involved somewhere) "learns" that an arbitrary term is heinous eand indirectly bans any discussion of that topic.

It's not like this unique to LLMs either. By some little trolling on internet you easily can turn hand "OK gesture" into a hate symbol of white supermacy. And fools will fall for it.

Re: I extracted the safety filters from Apple Intelligence models

#199
No shoot, bombs or bombers? I guess apple isnt interested in military contracts. Or, frankly, any work for world peace organizations dedicated to detecting and preventing genocide. And without talk of losing lives, much of the gaming industry is out too.

But i dont see the really bad stuff, the stuff i wont even type here. I guess that remains fair game. Apple's priorities remain as weird as ever.

Re: I extracted the safety filters from Apple Intelligence models

#200
post #111

Earlier quoted context omitted.

It's totally performative. There's no way to stay ahead of the new language that people create. At what point do the new words become the actual words? Are there many instances of people using unalive IRL?

> There's no way to stay ahead of the new language that people create. I'm imagining a new exploit: After someone says something totally innocent, people gang up in the comments to act like a terrible vicious slur has been said, and then the moderation system (with an LLM involved somewhere) "learns" that an arbitrary term is heinous eand indirectly bans any discussion of that topic.

The first half of that already happened with the OK gesture: https://www.bbc.co.uk/news/newsbeat-49837898.

Though it would be fun to see what happens if an LLM if used to ban anything that tends to generate heated exchanges. It would presumably learn to ban racial terms, politics and politicians and words like "immigrant" (i.e. basically the list in this repo), but what else could it be persuaded to ban? Vim and Emacs? SystemD? Anything involving cyclists? Parenting advice?

Post reply on HN