Live data from Hacker News

I extracted the safety filters from Apple Intelligence models

github.com

251–260 of 455 posts

Re: I extracted the safety filters from Apple Intelligence models

#251
post #215

Earlier quoted context omitted.

This is the rhetorical tactic of false equivalence. State censorship by an autocracy with the objective of population control is not the same thing as a private company inside a democracy censoring their product to avoid bad press and maintain goodwill for shareholders. If you want solid proof that it's not the same thing, see all the uncensored open weights models that you can freely download and use without fear of…

> is not the same thing as a private company inside a democracy censoring their product to avoid bad press and Yet this private company has more power and influence than most countries. And there are several such companies. We already live in sci fi corporate dystopia, we just haven't fully realised it yet.

People think a trillion dollar brainwashing industry is absolutely fine because of “democracy”, completely ignoring that all you have to do is use a century of experience convincing people to act against their own interests can deliver whatever you want.

Often the same people who think America is fine and safe are the ones who whine about the “main stream media” and “sheeple”.

Re: I extracted the safety filters from Apple Intelligence models

#252

Earlier quoted context omitted.

The first half of that already happened with the OK gesture: https://www.bbc.co.uk/news/newsbeat-49837898 . Though it would be fun to see what happens if an LLM if used to ban anything that tends to generate heated exchanges. It would presumably learn to ban racial terms, politics and politicians and words like "immigrant" (i.e. basically the list in this repo), but what else could it be persuaded to ban? Vim and Ema…

People weren't using the OK gesture innocently. After 4chan trolls decided to start pretending it was a white supremacist symbol, actual white supremacists started using it as a symbol.

All 10 of them?

What about the other 7-8 billion people still using it normally?

Re: I extracted the safety filters from Apple Intelligence models

#253

Some of the combinations are a bit weird, This one has lots of stuff avoiding death....together with a set ensuring all the Apple brands have the correct capitalisation. Priorities hey! https://github.com/BlueFalconHD/apple_generative_model_safet...

So it blocks it from suggesting to "execute" a file or "pass on" some information.

Yahoo had this problem years ago when they rewrote emails to avoid the term "eval". (trying to filter dangerous javascript) Famously producing the word "medireview".

Re: I extracted the safety filters from Apple Intelligence models

#254
post #107
post #43

Earlier quoted context omitted.

It would also substantially disrupt the generation process: a model which sees B0ris and not Boris is going to struggle to actually associate that input to the politician since it won't be well represented in the training set (and on the output side the same: if it does make the association, a reasoning model for example would include the proper name in the output first at which point the supervisor process can rejec…

"Draw a picture of a gorgon with the face of the 2024 Prime Minister of UK."

There were two.

Re: I extracted the safety filters from Apple Intelligence models

#255
post #198
post #111

Earlier quoted context omitted.

> There's no way to stay ahead of the new language that people create. I'm imagining a new exploit: After someone says something totally innocent, people gang up in the comments to act like a terrible vicious slur has been said, and then the moderation system (with an LLM involved somewhere) "learns" that an arbitrary term is heinous eand indirectly bans any discussion of that topic.

It's not like this unique to LLMs either. By some little trolling on internet you easily can turn hand "OK gesture" into a hate symbol of white supermacy. And fools will fall for it.

It's hack journalists reporting on BS totally fringe activity as if it's "a thing", and then idiots who take their cues from them

Re: I extracted the safety filters from Apple Intelligence models

#257
post #198

Earlier quoted context omitted.

It's not like this unique to LLMs either. By some little trolling on internet you easily can turn hand "OK gesture" into a hate symbol of white supermacy. And fools will fall for it.

...and then the bigots will fall for it too, and start using it in earnest, completing the cycle.

who cares what the bigots use?

If the bigots start using "thank you" as some code word, should we stop saying it, lest we pollute our non-bigoted discussions?

bigots drink coffee too, maybe we should stop drinking it, because something-something...

Re: I extracted the safety filters from Apple Intelligence models

#258

Some of the combinations are a bit weird, This one has lots of stuff avoiding death....together with a set ensuring all the Apple brands have the correct capitalisation. Priorities hey! https://github.com/BlueFalconHD/apple_generative_model_safet...

Interesting that it didn't seem to include "unalive". Which as a phenomenon is so very telling that no one actually cares what people are really saying. Everyone, including the platforms knows what that means. It's all performative.

> Everyone, including the platforms knows what that means.

Well, that's what happens when you let an enemy nation control one of the most biggest social networks there is. They just go try and see how far they can go.

On the other hand, Americans and their fear of four letter words or, gasp, exposed nipples are just as braindead.

Re: I extracted the safety filters from Apple Intelligence models

#259

Earlier quoted context omitted.

The problem with blocking names of politicians: the list of “notable politicians” is not only highly country-specific, it is also constantly changing-someone who is a near nobody today in a few more years could be a major world leader (witness the phenomenal rise of Barack Obama from yet another state senator in 2004-there’s close to 2000 of them-to US President 5 years later.) Will they put in the ongoing effort to…

Fun fact: There was at least on dip in Berkshire Hathaway stock, when Anne Hathaway got sick

Even if your keyword searching trading bot is smart enough to know it's unrelated, knowing there's dumber bots out there is information you can base trades on.

Re: I extracted the safety filters from Apple Intelligence models

#260

Earlier quoted context omitted.

Well that’s sad. They can’t even face the word ?

It’s not about whether they can face it. The younger generations are more in tune with mental health and topics like suicide than any previous generation. The etymology of the euphemism was about avoiding online censorship, while its “IRL” usage was merely absorbed through familiarity from the online usage.

>more in tune with mental health and topics like suicide than any previous generation.

More in such a fad than any previous generation

Post reply on HN