Live data from Hacker News

I extracted the safety filters from Apple Intelligence models

github.com

111–120 of 455 posts

Re: I extracted the safety filters from Apple Intelligence models

#111

Earlier quoted context omitted.

Interesting that it didn't seem to include "unalive". Which as a phenomenon is so very telling that no one actually cares what people are really saying. Everyone, including the platforms knows what that means. It's all performative.

It's totally performative. There's no way to stay ahead of the new language that people create. At what point do the new words become the actual words? Are there many instances of people using unalive IRL?

> There's no way to stay ahead of the new language that people create.

I'm imagining a new exploit: After someone says something totally innocent, people gang up in the comments to act like a terrible vicious slur has been said, and then the moderation system (with an LLM involved somewhere) "learns" that an arbitrary term is heinous eand indirectly bans any discussion of that topic.

Re: I extracted the safety filters from Apple Intelligence models

#112

Wow, this is pretty silly. If things are like this at Apple I’m not sure what to think. https://github.com/BlueFalconHD/apple_generative_model_safet... EDIT: just to be clear, things like this are easily bypassed. “Boris Johnson”=>”B0ris Johnson” will skip right over the regex and will be recognized just fine by an LLM.

Sounds like UK politics is taboo?

All politics is taboo, except the sort that helps Apple get richer. (Or any other company, in that company's "safety" filters)

Re: I extracted the safety filters from Apple Intelligence models

#114

Earlier quoted context omitted.

Interesting that it didn't seem to include "unalive". Which as a phenomenon is so very telling that no one actually cares what people are really saying. Everyone, including the platforms knows what that means. It's all performative.

It's totally performative. There's no way to stay ahead of the new language that people create. At what point do the new words become the actual words? Are there many instances of people using unalive IRL?

A specialized AI could do it as well as any human.

The future will be AIs all the way down...

Re: I extracted the safety filters from Apple Intelligence models

#116
post #97

Earlier quoted context omitted.

Yes, proper nouns are capitalized. And of course it's much worse for a company's published works to not respect branding-- a trademark only exists if it is actively defended. Official marketing material by a company has been used as legal evidence that their trademark has been genericized: >In one example, the Otis Elevator Company's trademark of the word "escalator" was cancelled following a petition from Toledo-bas…

Using a trademark as a noun is automatically genericizing. Capitalization of a noun is irrelevant to trademark. Even Apple corporation says that in their trademark guidance page, despite constantly breaking their own rule, when they call through iPhone phones "iPhone". But Apple, like founder Steve Jobs, believes the rules don't apply to them. https://www.apple.com/legal/intellectual-property/trademark/...

That explains why Steve Jobs never said “buy an iPhone” or “buy the iPhone” but “buy iPhone” (They always use it without “the” or “a”, like “buying a brand”).

Re: I extracted the safety filters from Apple Intelligence models

#117
post #3

There’s got to be a way to turn these lists of “naughty words” into shibboleths somehow.

Like asking sensitive employment candidates about Kim Jong Un's roundness to check if they're North Korean spies, we could ask humans what they think about Trump and Palestine to check if they're computers.

However, I think about half of real humans would also fail the test.

Re: I extracted the safety filters from Apple Intelligence models

#118

Earlier quoted context omitted.

It's totally performative. There's no way to stay ahead of the new language that people create. At what point do the new words become the actual words? Are there many instances of people using unalive IRL?

This question is sort of the same as asking why the universal translator wasn't able to translate the metaphor language of the Star Trek episode Darmok. Surely if the metaphor has become the first order meaning then there's no litteral meaning anymore.

The only reason kids started using "unalive" is to get around Youtube filters that disallow the use of the word "kill"

Re: I extracted the safety filters from Apple Intelligence models

#119
post #84

Earlier quoted context omitted.

Apple's 1984 ad is so hypocritical today. This is Apple actively steering public thought. No code - anywhere - should look like this. I don't care if the politicians are right, left, or authoritarian. This is wrong.

Why is this wrong? Applying special treatment to politically exposed persons has been standard practice in every high risk industry for a very long time. The simple fact is that people get extremely emotional about politicians, politicians both receive obscene amounts of abuse, and have repeatedly demonstrated they’re not above weaponising tools like this for their own goals. Seems perfectly reasonable that Apple doe…

I can Google for any of these people, and I can get real results with real information.
Post reply on HN