Live data from Hacker News

I extracted the safety filters from Apple Intelligence models

github.com

21–30 of 455 posts

Re: I extracted the safety filters from Apple Intelligence models

#21
post #17

I’m going to change my name to “Granular Mango Serpent” just to see what those keywords are for in their safety instructions.

Granular Mango Serpent is the new David Meyer.

https://arstechnica.com/information-technology/2024/12/certa...

Re: I extracted the safety filters from Apple Intelligence models

#22

Wow, this is pretty silly. If things are like this at Apple I’m not sure what to think. https://github.com/BlueFalconHD/apple_generative_model_safet... EDIT: just to be clear, things like this are easily bypassed. “Boris Johnson”=>”B0ris Johnson” will skip right over the regex and will be recognized just fine by an LLM.

It's not silly. I would bet 99% of the users don't care that much to do that. A hardcoded regex like this is a good first layer/filter, and very efficient

Re: I extracted the safety filters from Apple Intelligence models

#23
post #12

Earlier quoted context omitted.

> Apple brands have the correct capitalisation. Priorities hey! To me that's really embarrassing and insecure. But I'm sure for branding people it's very important.

Legal requirement to maintain a trademark.

In what way would (A|a)pple's own AI writing "imac" endanger the trademark? Is capitalisation even part of a word-based trademark?

I'm more surprised they don't have a rule to do that rather grating s/the iPhone/iPhone/ transform (or maybe it's in a different file?).

Re: I extracted the safety filters from Apple Intelligence models

#24

Wow, this is pretty silly. If things are like this at Apple I’m not sure what to think. https://github.com/BlueFalconHD/apple_generative_model_safet... EDIT: just to be clear, things like this are easily bypassed. “Boris Johnson”=>”B0ris Johnson” will skip right over the regex and will be recognized just fine by an LLM.

Sounds like UK politics is taboo?

Re: I extracted the safety filters from Apple Intelligence models

#25

Some of the combinations are a bit weird, This one has lots of stuff avoiding death....together with a set ensuring all the Apple brands have the correct capitalisation. Priorities hey! https://github.com/BlueFalconHD/apple_generative_model_safet...

Interesting that it didn't seem to include "unalive". Which as a phenomenon is so very telling that no one actually cares what people are really saying. Everyone, including the platforms knows what that means. It's all performative.

It's totally performative. There's no way to stay ahead of the new language that people create.

At what point do the new words become the actual words? Are there many instances of people using unalive IRL?

Re: I extracted the safety filters from Apple Intelligence models

#26

Earlier quoted context omitted.

Interesting that it didn't seem to include "unalive". Which as a phenomenon is so very telling that no one actually cares what people are really saying. Everyone, including the platforms knows what that means. It's all performative.

It's totally performative. There's no way to stay ahead of the new language that people create. At what point do the new words become the actual words? Are there many instances of people using unalive IRL?

It depends on if you think that something is less real because it’s transmitted digitally.

Re: I extracted the safety filters from Apple Intelligence models

#27

Wow, this is pretty silly. If things are like this at Apple I’m not sure what to think. https://github.com/BlueFalconHD/apple_generative_model_safet... EDIT: just to be clear, things like this are easily bypassed. “Boris Johnson”=>”B0ris Johnson” will skip right over the regex and will be recognized just fine by an LLM.

I doubt the purpose here is so much to prevent someone from intentionally side stepping the block. It's more likely here to avoid the sort of headlines you would expect to see if someone was suggested "I wish ${politician} would die" as a response to an email mentioning that politician. In general you should view these sorts of broad word filters as looking to short circuit the "think of the children" reactions to Tiny Tim's phone suggesting not that God should "bless us, every one", but that God should "kill us, every one". A dumb filter like this is more than enough for that sort of thing.

Re: I extracted the safety filters from Apple Intelligence models

#29

Are you sure it's fully deobfuscated? What's up with reject phrases like "Granular mango serpent"?

the one at the bottom of the README spells out xcode wyvern illustrous laments darkness

read every good expletive “xxx”
Post reply on HN