Live data from Hacker News

I extracted the safety filters from Apple Intelligence models

github.com

151–160 of 455 posts

Re: I extracted the safety filters from Apple Intelligence models

#151
post #43
post #27

Earlier quoted context omitted.

I doubt the purpose here is so much to prevent someone from intentionally side stepping the block. It's more likely here to avoid the sort of headlines you would expect to see if someone was suggested "I wish ${politician} would die" as a response to an email mentioning that politician. In general you should view these sorts of broad word filters as looking to short circuit the "think of the children" reactions to Ti…

It would also substantially disrupt the generation process: a model which sees B0ris and not Boris is going to struggle to actually associate that input to the politician since it won't be well represented in the training set (and on the output side the same: if it does make the association, a reasoning model for example would include the proper name in the output first at which point the supervisor process can rejec…

No it doesn't disrupt. This is a well known capability of LLMs. Most models don't even point out a mistake they just carry on.

https://chatgpt.com/share/686b1092-4974-8010-9c33-86036c88e7...

Re: I extracted the safety filters from Apple Intelligence models

#152
post #37

Earlier quoted context omitted.

This is just policy and alignment from Apple. Just because the Internet says a bunch of junk doesn't mean you want your model spewing it.

sure but models also can't see any truth on their own. They are literally butchered and lobotomized with filters and such. Even high IQ people struggle with certain truth after reading a lot, how is these models going to find it with so much filters?

> sure but models also can't see any truth on their own. They are literally butchered and lobotomized with filters and such.

The one is unrelated to the other.

> Even high IQ people struggle with certain truth after reading a lot,

Huh?

Re: I extracted the safety filters from Apple Intelligence models

#153
post #111

Earlier quoted context omitted.

> There's no way to stay ahead of the new language that people create. I'm imagining a new exploit: After someone says something totally innocent, people gang up in the comments to act like a terrible vicious slur has been said, and then the moderation system (with an LLM involved somewhere) "learns" that an arbitrary term is heinous eand indirectly bans any discussion of that topic.

I'm pretty sure this can work human moderators rather than an LLM, too.

Most of the human moderators hired by OpenAI to train LLMs, many of them based in Africa and South America, were exposed to disturbing content and have been deeply affected by it.

Karen Hao interviewed many of them in her latest bestselling book, which explores the human cost behind the OpenAI boom:

https://www.goodreads.com/book/show/222725518-empire-of-ai

Re: I extracted the safety filters from Apple Intelligence models

#154

Is this related in any way to Core ML model encryption ( https://developer.apple.com/documentation/coreml/encrypting-... )? I find that feature a little bizarre because Apple has historically avoided providing any kind of DRM solution for app asset protection.

Nope. This is a separate system. It’s not even abstracted for any asset, it is specifically only for these overrides. The decryption is done in the ModelCatalog private framework.

Re: I extracted the safety filters from Apple Intelligence models

#155

China calls it "harmonious society", we call it "safety". Censorship by any other name would be just as effective for manipulating the thoughts of the populace. It's not often that you get to see stuff like this.

I don't think it's controversial or unsurprising at all that a company doesn't want their random sentence generator to spit out 'brand damaging' sentences. You know the field day media would have Apple's new feature summarises a text message as "Jane thinks Anthony Albanese should die".

If that's what the message actually said, why would the media be complaining? Or do you mean false positives?

Re: I extracted the safety filters from Apple Intelligence models

#156
post #12

Earlier quoted context omitted.

> Apple brands have the correct capitalisation. Priorities hey! To me that's really embarrassing and insecure. But I'm sure for branding people it's very important.

Legal requirement to maintain a trademark.

In their own marketing language, sure, but to force this on their users' speech?

Consider that these models, among other things, power features such as "proofread" or "rewrite professionally".

Re: I extracted the safety filters from Apple Intelligence models

#157

Earlier quoted context omitted.

Interesting that it didn't seem to include "unalive". Which as a phenomenon is so very telling that no one actually cares what people are really saying. Everyone, including the platforms knows what that means. It's all performative.

yo, these are businesses. It's not performative, its CYA. They care because of legal reasons, not moral or ethical.

Seriously. I feel like “performative” gets applied to anything imperfect. They’ll never stop 100% of murders, so these laws against it are just performative…

Re: I extracted the safety filters from Apple Intelligence models

#158

Earlier quoted context omitted.

Interesting that it didn't seem to include "unalive". Which as a phenomenon is so very telling that no one actually cares what people are really saying. Everyone, including the platforms knows what that means. It's all performative.

yo, these are businesses. It's not performative, its CYA. They care because of legal reasons, not moral or ethical.

Does adding a trivial word filter even make any sense from a legal point of view, especially when this one seems to be filtering out words describing concepts that can be pretty easily paraphrased?

A regex sounds like a bad solution for profanity, but like an even worse one to bolt onto a thing that's literally designed to be able to communicate like a human and could probably easily talk its way around guardrails if it were so inclined.

Re: I extracted the safety filters from Apple Intelligence models

#159

Earlier quoted context omitted.

In what way would (A|a)pple's own AI writing "imac" endanger the trademark? Is capitalisation even part of a word-based trademark? I'm more surprised they don't have a rule to do that rather grating s/the iPhone/iPhone/ transform (or maybe it's in a different file?).

Yes, proper nouns are capitalized. And of course it's much worse for a company's published works to not respect branding-- a trademark only exists if it is actively defended. Official marketing material by a company has been used as legal evidence that their trademark has been genericized: >In one example, the Otis Elevator Company's trademark of the word "escalator" was cancelled following a petition from Toledo-bas…

Sure, but software that autocompletes/rewords users' emails and text messages is not marketing material.

Otherwise, why stop there? Why not have the macOS keyboard driver or Safari prevent me from typing "Iphone"? Why not have iOS edit my voice if I call their Bluetooth headphones "earbuds pro" in a phone call?

Re: I extracted the safety filters from Apple Intelligence models

#160
post #111

Earlier quoted context omitted.

> There's no way to stay ahead of the new language that people create. I'm imagining a new exploit: After someone says something totally innocent, people gang up in the comments to act like a terrible vicious slur has been said, and then the moderation system (with an LLM involved somewhere) "learns" that an arbitrary term is heinous eand indirectly bans any discussion of that topic.

Hey I was pro-skub waaaay before all the anti-skub people switched sides.

Skub is a real slur tho so that one doesn’t work
Post reply on HN