Live data from Hacker News

I extracted the safety filters from Apple Intelligence models

github.com

301–310 of 455 posts

Re: I extracted the safety filters from Apple Intelligence models

#301

Earlier quoted context omitted.

The first half of that already happened with the OK gesture: https://www.bbc.co.uk/news/newsbeat-49837898 . Though it would be fun to see what happens if an LLM if used to ban anything that tends to generate heated exchanges. It would presumably learn to ban racial terms, politics and politicians and words like "immigrant" (i.e. basically the list in this repo), but what else could it be persuaded to ban? Vim and Ema…

The OK gesture has always been very inappropriate in most parts of the world.

> The OK gesture has always been very inappropriate in most parts of the world.

No, it isn't, and especially hasn't been historically. The negative connotations are overwhelmingly modern.

The areas where it is very inappropriate right now tally up to maybe 1 billion people*. That's pretty far from "most". For everyone else it is mostly positive, neutral, or meaningless.

*Brazil, Turkey, Iran, Iraq, Saudi Arabia, Greece, Italy, Spain, Russia, Ukraine, Belarus, other parts of Eastern Europe

Re: I extracted the safety filters from Apple Intelligence models

#302

Earlier quoted context omitted.

...and then the bigots will fall for it too, and start using it in earnest, completing the cycle.

who cares what the bigots use? If the bigots start using "thank you" as some code word, should we stop saying it, lest we pollute our non-bigoted discussions? bigots drink coffee too, maybe we should stop drinking it, because something-something...

I don’t think we should treat human interactions like a technical problem, where we look for edge cases and outlandish hypotheticals to probe the edges of what is possible.

If “thank you” became widely associated with bigots, and had some negative meaning, to the point where it genuinely distressed people, I’d avoid it. I think it has a widespread enough normal meaning that there’s almost no chance of that happening, but it isn’t impossible.

Re: I extracted the safety filters from Apple Intelligence models

#303

Earlier quoted context omitted.

It's totally performative. There's no way to stay ahead of the new language that people create. At what point do the new words become the actual words? Are there many instances of people using unalive IRL?

> Are there many instances of people using unalive IRL As a parent of a teenager, I see them use "unalive" non-ironically as a synonym for "suicide" in all contexts, including IRL.

“Unalive” is sort of… awkward in that silly online way. But, we also have phrase like “off oneself,” or just euphemistically describing the person as having died. It’s always been a difficult topic to talk about, I don’t understand using it as a specific example of gen-Z fragility.

Just that they suck at coming up with pithy new slang terms.

Re: I extracted the safety filters from Apple Intelligence models

#304
post #296

Earlier quoted context omitted.

After the big surprise of seeing at work a list with all my personal purchases included in a big set of documents to which I, together with a great number of other colleagues, had access, I went immediately to the bank and I reported the fact. After some days had passed without seeing any consequence, I went again, this time discussing with some supervising employee, who attempted to convince me that this is some kin…

Behavior isn't what needs to change here. It's a poor system design. Humans make mistakes. Systems prevent mistakes. Do you think the mistake would have happened if a machine checked the numbers vs the address? How about if a 2nd person looked it over? How about both? In this case a computer could have easily flagged an address mismatch between your account number and the receiver (your work).

Thank you, that's what I intended to say.

Re: I extracted the safety filters from Apple Intelligence models

#306
post #274

Some of the combinations are a bit weird, This one has lots of stuff avoiding death....together with a set ensuring all the Apple brands have the correct capitalisation. Priorities hey! https://github.com/BlueFalconHD/apple_generative_model_safet...

Also feels like some of these would match totally innocuous usage. "I'm overloaded for work, I'd be happy if you took some of it off me." "The client seems to have passed on the proposed changes." Both of those would match the "death regexes". Seems we haven't learned from the "glbutt of wine" problem of content filtering even decades later - the learnings of which are that you simply cannot do content filtering base…

Aka the 'Scunthorpe Problem'

Re: I extracted the safety filters from Apple Intelligence models

#307
post #235

Earlier quoted context omitted.

> I considered that an appropriate punishment would have been a pay cut for a few months This can absolutely cripple a family, I'd be really cautious wishing that upon someone if they wronged you without malice, though I completely understand where you are coming from. In this case at the very least, I'd want to know what went wrong and what they’re doing to make sure it doesn’t happen again. From a software-engineer…

After the big surprise of seeing at work a list with all my personal purchases included in a big set of documents to which I, together with a great number of other colleagues, had access, I went immediately to the bank and I reported the fact. After some days had passed without seeing any consequence, I went again, this time discussing with some supervising employee, who attempted to convince me that this is some kin…

Thanks for sharing. Sounds like they have (hopefully _had_) a really messy system in place.

And just to be clear, I didn’t mean to downplay what happened to you, I completely understand how serious it is.

Re: I extracted the safety filters from Apple Intelligence models

#309
post #308

Nice to see that we are protected from talking about these weird old dolls: https://en.wikipedia.org/wiki/Golliwog https://github.com/BlueFalconHD/apple_generative_model_safet...

Well, they're not only weird, they're obviously racist doll.

Re: I extracted the safety filters from Apple Intelligence models

#310

Earlier quoted context omitted.

...and then the bigots will fall for it too, and start using it in earnest, completing the cycle.

who cares what the bigots use? If the bigots start using "thank you" as some code word, should we stop saying it, lest we pollute our non-bigoted discussions? bigots drink coffee too, maybe we should stop drinking it, because something-something...

>who cares what the bigots use

you'd think so, but people often operate where multiple contexts could be valid.

Just as a thought experiment, if the eggplant emoji was used to denote "ok" in messaging and then people starting appropriating it for a sexual context, would you or the general public think twice about continuing to use it to mean "ok" on the off chance the other side may misinterpret the meaning?

I would say most likely yes.

Post reply on HN