Live data from Hacker News

I extracted the safety filters from Apple Intelligence models

github.com

291–300 of 455 posts

Re: I extracted the safety filters from Apple Intelligence models

#291

Earlier quoted context omitted.

Already happened: "We found instances of the model attempting to write self-propagating worms, fabricating legal documentation, and leaving hidden notes to future instances of itself all in an effort to undermine its developers' intentions" [1]. [1] https://www.axios.com/2025/05/23/anthropic-ai-deception-risk

Note that all these things are in the training data. That's all that is. I'm trying to remember which movie it was where a man left notes to himself because he had memory loss, as I never saw that movie. That's the sort of thing where an AI could easily tell me with very little back-and-forth and be correct, because it's broadly popular information that's in the training data and just I don't remember it. By the same…

I think you mean the movie "Memento"

Re: I extracted the safety filters from Apple Intelligence models

#292

Earlier quoted context omitted.

The first half of that already happened with the OK gesture: https://www.bbc.co.uk/news/newsbeat-49837898 . Though it would be fun to see what happens if an LLM if used to ban anything that tends to generate heated exchanges. It would presumably learn to ban racial terms, politics and politicians and words like "immigrant" (i.e. basically the list in this repo), but what else could it be persuaded to ban? Vim and Ema…

People weren't using the OK gesture innocently. After 4chan trolls decided to start pretending it was a white supremacist symbol, actual white supremacists started using it as a symbol.

then congratulations on making white supremacists define your langyage

Re: I extracted the safety filters from Apple Intelligence models

#293
post #111

Earlier quoted context omitted.

> There's no way to stay ahead of the new language that people create. I'm imagining a new exploit: After someone says something totally innocent, people gang up in the comments to act like a terrible vicious slur has been said, and then the moderation system (with an LLM involved somewhere) "learns" that an arbitrary term is heinous eand indirectly bans any discussion of that topic.

The first half of that already happened with the OK gesture: https://www.bbc.co.uk/news/newsbeat-49837898 . Though it would be fun to see what happens if an LLM if used to ban anything that tends to generate heated exchanges. It would presumably learn to ban racial terms, politics and politicians and words like "immigrant" (i.e. basically the list in this repo), but what else could it be persuaded to ban? Vim and Ema…

The OK gesture has always been very inappropriate in most parts of the world.

Re: I extracted the safety filters from Apple Intelligence models

#294

Earlier quoted context omitted.

It's totally performative. There's no way to stay ahead of the new language that people create. At what point do the new words become the actual words? Are there many instances of people using unalive IRL?

> Are there many instances of people using unalive IRL? In my experience yes. This is already commonplace. Mostly, but not exclusively, amongst the younger generation.

I think it stemmed from content creators using it to avoid platform filters (even if video is not removed it gets deprioritized, at least on YT) and kids repeat it

Re: I extracted the safety filters from Apple Intelligence models

#295

Earlier quoted context omitted.

Interesting that it didn't seem to include "unalive". Which as a phenomenon is so very telling that no one actually cares what people are really saying. Everyone, including the platforms knows what that means. It's all performative.

It's totally performative. There's no way to stay ahead of the new language that people create. At what point do the new words become the actual words? Are there many instances of people using unalive IRL?

Lucky developers who wrote these rules live in totality different world at far distance from people

Re: I extracted the safety filters from Apple Intelligence models

#296
post #235

Earlier quoted context omitted.

> I considered that an appropriate punishment would have been a pay cut for a few months This can absolutely cripple a family, I'd be really cautious wishing that upon someone if they wronged you without malice, though I completely understand where you are coming from. In this case at the very least, I'd want to know what went wrong and what they’re doing to make sure it doesn’t happen again. From a software-engineer…

After the big surprise of seeing at work a list with all my personal purchases included in a big set of documents to which I, together with a great number of other colleagues, had access, I went immediately to the bank and I reported the fact. After some days had passed without seeing any consequence, I went again, this time discussing with some supervising employee, who attempted to convince me that this is some kin…

Behavior isn't what needs to change here. It's a poor system design. Humans make mistakes. Systems prevent mistakes.

Do you think the mistake would have happened if a machine checked the numbers vs the address? How about if a 2nd person looked it over? How about both?

In this case a computer could have easily flagged an address mismatch between your account number and the receiver (your work).

Re: I extracted the safety filters from Apple Intelligence models

#297
post #158

Earlier quoted context omitted.

yo, these are businesses. It's not performative, its CYA. They care because of legal reasons, not moral or ethical.

Does adding a trivial word filter even make any sense from a legal point of view, especially when this one seems to be filtering out words describing concepts that can be pretty easily paraphrased? A regex sounds like a bad solution for profanity, but like an even worse one to bolt onto a thing that's literally designed to be able to communicate like a human and could probably easily talk its way around guardrails if…

To a lawyer? Yes. I'm pretty sure a lawyer can easily search through all the business law and "Trivially" find case laws connected to words.

We're not talking about logical inference, we're talking about CYA.

Re: I extracted the safety filters from Apple Intelligence models

#298

Earlier quoted context omitted.

Interesting that it didn't seem to include "unalive". Which as a phenomenon is so very telling that no one actually cares what people are really saying. Everyone, including the platforms knows what that means. It's all performative.

It's also a shining example of American puritanism. Asian models or those in Europe are far less censored.

I'm sure this has more to do with legal liability than morals.

Re: I extracted the safety filters from Apple Intelligence models

#299
post #111

Earlier quoted context omitted.

> There's no way to stay ahead of the new language that people create. I'm imagining a new exploit: After someone says something totally innocent, people gang up in the comments to act like a terrible vicious slur has been said, and then the moderation system (with an LLM involved somewhere) "learns" that an arbitrary term is heinous eand indirectly bans any discussion of that topic.

The first half of that already happened with the OK gesture: https://www.bbc.co.uk/news/newsbeat-49837898 . Though it would be fun to see what happens if an LLM if used to ban anything that tends to generate heated exchanges. It would presumably learn to ban racial terms, politics and politicians and words like "immigrant" (i.e. basically the list in this repo), but what else could it be persuaded to ban? Vim and Ema…

It would probably ban discussion of censorship.

Re: I extracted the safety filters from Apple Intelligence models

#300
post #157

Earlier quoted context omitted.

Seriously. I feel like “performative” gets applied to anything imperfect. They’ll never stop 100% of murders, so these laws against it are just performative…

It seems more like banning specifically stabbing, shooting, strangulation and blunt impact rather then murder in general, and then just allowing killing by pushing out of windows because people figured out that it's not covered by existing laws. But no one important seems to be kicking up a fuss right now, so well allow it, as the lack of fuss is the key thing thing here. Not that I think going on a thorough mission…

The point is: "perfomative" refers to aping Ethical and Moral behaviors. That is _not_ why Apple would do this. They would do this because Legally, they could be culpable if an LLM told a 14 year old to do _anything_ thats illegal.

That's all. I'm constantly amazed how this basic CYA legal world escapes into griping about social culture war nonsense.

Post reply on HN