Live data from Hacker News

I extracted the safety filters from Apple Intelligence models

github.com

431–440 of 455 posts

Re: I extracted the safety filters from Apple Intelligence models

#431
post #159

Earlier quoted context omitted.

Sure, but software that autocompletes/rewords users' emails and text messages is not marketing material. Otherwise, why stop there? Why not have the macOS keyboard driver or Safari prevent me from typing "Iphone"? Why not have iOS edit my voice if I call their Bluetooth headphones "earbuds pro" in a phone call?

Because in regards to the rights to a trademark, what is critical is the use of the word in trade -- not just "marketing material" nor your phone calls to your friends.

So if I write a business email to my colleague, if Apple doesn't autocorrect "Iphone" to "iPhone" in it, they risk losing the trademark?

Re: I extracted the safety filters from Apple Intelligence models

#432

Earlier quoted context omitted.

then congratulations on making white supremacists define your langyage

Do you still use swastikas as symbols of peace and love because you don't want white supremacists to define your language? I strongly doubt you do that. Whether you like it or not, the Nazis defined what the swastika means now.

No, because western culture never really did. However the countries who have been using it for at least thousands of years in Buddhism are still using it just fine.

In fact there was a recent thing with one of the BTS members' uniform (worn during mandatory military service period in South Korea), which had the regular (not tilted) swastika on it because he was assigned to religious duties.

And of course the western world/media ran away with it. Plenty of absolutely brain dead people out there who couldn't research a topic to gain an understanding to save their lives.

Re: I extracted the safety filters from Apple Intelligence models

#433

Earlier quoted context omitted.

I don’t think we should treat human interactions like a technical problem, where we look for edge cases and outlandish hypotheticals to probe the edges of what is possible. If “thank you” became widely associated with bigots, and had some negative meaning, to the point where it genuinely distressed people, I’d avoid it. I think it has a widespread enough normal meaning that there’s almost no chance of that happening,…

This approach gives people you vehemently disagree with a lot of power over you.

Yup, it's basically saying "I'll let bullies win".

Re: I extracted the safety filters from Apple Intelligence models

#434
post #174

Earlier quoted context omitted.

It’s not about whether they can face it. The younger generations are more in tune with mental health and topics like suicide than any previous generation. The etymology of the euphemism was about avoiding online censorship, while its “IRL” usage was merely absorbed through familiarity from the online usage.

But unalive self is suicide and unalive is just death, right? For example, You can unalive other people against their will...

'An hero' came before it but that was as a meme.

Unalive is mostly to avoid censorship same as ahh. But once they enter common usage it's not really about censorship anymore.

Re: I extracted the safety filters from Apple Intelligence models

#435

Earlier quoted context omitted.

The first half of that already happened with the OK gesture: https://www.bbc.co.uk/news/newsbeat-49837898 . Though it would be fun to see what happens if an LLM if used to ban anything that tends to generate heated exchanges. It would presumably learn to ban racial terms, politics and politicians and words like "immigrant" (i.e. basically the list in this repo), but what else could it be persuaded to ban? Vim and Ema…

The OK gesture has always been very inappropriate in most parts of the world.

The OK gesture has been the standard gesture for saying OK for scuba diving all over the world (PADI). I have used it all over the world on my scuba diving trips and have never had any problem or negative reaction to it.

Re: I extracted the safety filters from Apple Intelligence models

#436

Earlier quoted context omitted.

The OK gesture has always been very inappropriate in most parts of the world.

> The OK gesture has always been very inappropriate in most parts of the world. No, it isn't, and especially hasn't been historically. The negative connotations are overwhelmingly modern. The areas where it is very inappropriate right now tally up to maybe 1 billion people*. That's pretty far from "most". For everyone else it is mostly positive, neutral, or meaningless. *Brazil, Turkey, Iran, Iraq, Saudi Arabia, Gree…

I use it in Brazil scuba diving as it's the universal PADI hand gesture for asking (and responding) if someone is OK and never had any issues or negative reactions.

The PADI standard gestures are used and recognized all over the world to mean these things.

https://blog.padi.com/scuba-diving-hand-signals/

Re: I extracted the safety filters from Apple Intelligence models

#437

Some of the combinations are a bit weird, This one has lots of stuff avoiding death....together with a set ensuring all the Apple brands have the correct capitalisation. Priorities hey! https://github.com/BlueFalconHD/apple_generative_model_safet...

Interesting that it didn't seem to include "unalive". Which as a phenomenon is so very telling that no one actually cares what people are really saying. Everyone, including the platforms knows what that means. It's all performative.

No leetspeak filters either.

Re: I extracted the safety filters from Apple Intelligence models

#438
post #429

Earlier quoted context omitted.

Yeah, those topics are definitely censored on big platforms but I have the impression that it relies of manual reporting. At least reddit feels like that because what you can say depends on the subreddit - not just the mods but what kinds of people visit it and what they report. No idea about youtube, videos are definitely censored using some automated means but it's still possible to get around it. E.g. some gun you…

All of these platforms except perhaps Reddit are using LLMs (and other ML/AI) for censoring and automated anti-abuse. Including the LLM platforms themselves. Manual reporting is an adjunct/additional method, and goes into the training data set after whatever manual intervention occurs too.

Not to sound like I am rejecting the possibility but can you tell me how you got that information? I would be very helpful for convincing people in general to have something more concrete to go on that a random comment.

Re: I extracted the safety filters from Apple Intelligence models

#439
post #429

Earlier quoted context omitted.

All of these platforms except perhaps Reddit are using LLMs (and other ML/AI) for censoring and automated anti-abuse. Including the LLM platforms themselves. Manual reporting is an adjunct/additional method, and goes into the training data set after whatever manual intervention occurs too.

Not to sound like I am rejecting the possibility but can you tell me how you got that information? I would be very helpful for convincing people in general to have something more concrete to go on that a random comment.

I build those systems at a company that you definitely are aware of. I can’t discuss it further due to my NDA.

Feel free to ignore that any of this exists of course - it makes our lives easier. It’s a constant arms race regardless.

Re: I extracted the safety filters from Apple Intelligence models

#440
post #431

Earlier quoted context omitted.

Because in regards to the rights to a trademark, what is critical is the use of the word in trade -- not just "marketing material" nor your phone calls to your friends.

So if I write a business email to my colleague, if Apple doesn't autocorrect "Iphone" to "iPhone" in it, they risk losing the trademark?

Your emails aren't very relevant. But the way Apple's represents their product is.
Post reply on HN