Live data from Hacker News

I extracted the safety filters from Apple Intelligence models

github.com

421–430 of 455 posts

Re: I extracted the safety filters from Apple Intelligence models

#421

Earlier quoted context omitted.

I don't think it's controversial or unsurprising at all that a company doesn't want their random sentence generator to spit out 'brand damaging' sentences. You know the field day media would have Apple's new feature summarises a text message as "Jane thinks Anthony Albanese should die".

When the choice is between 1. "avoid tarnishing my own brand" and 2. "doing what the user requested," corporations will always choose option 1. Who is this software supposed to be serving, anyway? I'm surprised MS Office still allows me to type "Microsoft can go suck a dick" into a document and Apple's Pages app still allows me to type "Apple are hypocritical jerks." I wonder how long until that won't be the case...

But so often these tools are used in a way that the user didn't explicitly request, like summarising notifications, or generating slideshows from your photo library.

Re: I extracted the safety filters from Apple Intelligence models

#422
post #415

Earlier quoted context omitted.

Can you give examples? The closest I've seen is autodetection of certain topics related to death and suicide and subsequently promoting some kind of "help" hotline. A friend also said google allows an interview with a pedophile on youtube but penalizes it in search results so much that it's (almost?) impossible to find even when using the exact name. But of course, if a topic is shadowbanned, it's hard to find out ab…

Guns (specific elements). Drugs (manufacture). Sexual topics. Cursing (too much). Large swathes of political topics. Crypto. It’s flip-flopped on specifics numerous times over the years, but these policies are easy to find. From demonitization, channel bans (direct and shadow), and creator bans. We can of course argue until we’re blue in the face about correctness or not (most are not unreasonable by some societal de…

Yeah, those topics are definitely censored on big platforms but I have the impression that it relies of manual reporting.

At least reddit feels like that because what you can say depends on the subreddit - not just the mods but what kinds of people visit it and what they report.

No idea about youtube, videos are definitely censored using some automated means but it's still possible to get around it. E.g. some gun youtubers avoided saying full-auto by saying more-semi-auto. So i don't think they use very sophisticated models or they don't are yet. This kind of thing is obvious to a human and even LLMs generate responses which say it's a tongue-in-cheek to avoid censorship.

Comments are also generally less censored. After that health insurance CEO got punished for mass murder and repeated bodily harm with an extra-legal death penalty, many people were openly supporting it. I can say it here too and nobody will care. Even LLMs (both US and Chinese, except Claude because Claude is trained by eggshell-walking suckers) readily generate estimates of how many people he caused to die or suffer.

The internet would look very different if companies started using state of the art models to detect undesirable-to-them speech. But also people would fight back more so it might just be a case of boiling the frog slowly.

Re: I extracted the safety filters from Apple Intelligence models

#423

Earlier quoted context omitted.

There is a huge difference between an honest mistake by an employee, and clear employee misconduct. Punishing employees for making honest mistakes, where appropriate process should have prevented error, is a horrific way to handle mistakes like this. It would be equivalent to personally punishing engineers every time they deployed code that contained bugs. Nobody would ever think that’s an acceptable thing to do, why…

This was not a honest mistake. It was completely reckless behavior, even if the guilt was distributed both on the employee who has not checked whether the information sent to external parties is information to which access is permitted for them and on the employees who did not implement a system that would check automatically for such mistakes. Moreover, the attempt made by multiple bank employees to hide the inciden…

Has it occurred to you that personally punishing employees would just create further incentive to hide errors? You just create a culture of fear, where any attempt to acknowledge mistakes and learn from them is punished rather than rewarded.

I have no idea why you think inflicting financial penalties on employees would result in better outcomes. You only need to look at some highly avoidable transit disasters in Japan to understand why a model of punishment produces worse outcomes, not better.

https://en.m.wikipedia.org/wiki/Amagasaki_derailment

There is a reason we have regulators (or at least we do in the UK). I can assure you that if this had happened in the UK, and the complaint raised to the Financial Ombudsman (FOS), there would have been hefty financial punishment for the bank. If there were repeated infractions, the FCA would step in to investigate, and possibly personally punish C-suite leaders for failing to build the needed processes and culture to both prevent, and learn from mistakes like this.

And I’m not speaking about theory, I’m speaking from personal experience. I know exactly what it’s like to be on the pointy end of both the FOS and FCAs gaze. It’s not a comfortable position for any team in any bank, and even less comfortable for senior leaders.

Re: I extracted the safety filters from Apple Intelligence models

#424

Earlier quoted context omitted.

You would hope that search would be a politically safe space to operate. But politicians find a way to ruin everything for short term political gain. https://arstechnica.com/tech-policy/2018/12/republicans-in-c...

I would hope! But no one actually believes Google is politically neutral do they?

Evidence suggests they’re about as neutral as you could hope.

It’s not like Google search is some kind special tool used only by the elite. It’s pretty trivial for political scientists to pump queries into Google and measure the results. Which is exactly what many have done.

There’s been plenty of independent research into political bias of Google search results, and plenty of lawsuits that have gone fishing via discovery for internal evidence of bias. As yet, nobody has found a smoking gun, or any real evidence of search result bias (on a political axis, the same can be said for commercial gain).

There are many problems with Google, and Google search. Google as an org isn’t politically neutral (although I have no idea how they could be). But political bias in their results isn’t one of those problems.

Re: I extracted the safety filters from Apple Intelligence models

#425
post #401

Earlier quoted context omitted.

Really? What does DeepSeek say about Tiananmen Square? I'm not aware of any German models, but if you find one you should ask it what it thinks about Palestine. ( Qwen Mistral is French, but I have no idea what stuff would be censored in France)

> I have no idea what stuff would be censored in France Being French, what is the most likely to be censored relates to the Nazis. Holocaust denial is a crime for instance. Hate speech in general, including racism, antisemitism, homophobia, sexism, etc... is less tolerated than in countries like the US that have a more "free for all" view of free speech. We also have strong anti-defamation laws, that can also apply t…

[deleted]

Re: I extracted the safety filters from Apple Intelligence models

#426

Earlier quoted context omitted.

People weren't using the OK gesture innocently. After 4chan trolls decided to start pretending it was a white supremacist symbol, actual white supremacists started using it as a symbol.

All 10 of them? What about the other 7-8 billion people still using it normally?

I promise you the world contains more than 10 white supremacists and less than 7,000,000,000 non-white-supremacists who regularly use the OK sign.

Re: I extracted the safety filters from Apple Intelligence models

#427

Earlier quoted context omitted.

> doesn't mean people actually believe machines are human. They don't have to believe it's a human. I know a person who admitted to arguing with an LLM.

Which still does not demonstrate that they believe it has opinions. Natural language is how you interact with an LLM -- interactions will mimic human interaction, even for those who realize it is not sentient.

They were under the impression they could in fact change the AI's mind. So yes, they did believe it has an opinion. They believed it was sentient and able to think for itself. Do not underestimate peoples inability to distinguish between a very clever Markov chain and actual intelligence. The future is going to be ... interesting.

Re: I extracted the safety filters from Apple Intelligence models

#428

Earlier quoted context omitted.

The OK gesture has always been very inappropriate in most parts of the world.

> The OK gesture has always been very inappropriate in most parts of the world. No, it isn't, and especially hasn't been historically. The negative connotations are overwhelmingly modern. The areas where it is very inappropriate right now tally up to maybe 1 billion people*. That's pretty far from "most". For everyone else it is mostly positive, neutral, or meaningless. *Brazil, Turkey, Iran, Iraq, Saudi Arabia, Gree…

> Greece

It's perfectly OK in Greece.

Re: I extracted the safety filters from Apple Intelligence models

#429
post #415

Earlier quoted context omitted.

Guns (specific elements). Drugs (manufacture). Sexual topics. Cursing (too much). Large swathes of political topics. Crypto. It’s flip-flopped on specifics numerous times over the years, but these policies are easy to find. From demonitization, channel bans (direct and shadow), and creator bans. We can of course argue until we’re blue in the face about correctness or not (most are not unreasonable by some societal de…

Yeah, those topics are definitely censored on big platforms but I have the impression that it relies of manual reporting. At least reddit feels like that because what you can say depends on the subreddit - not just the mods but what kinds of people visit it and what they report. No idea about youtube, videos are definitely censored using some automated means but it's still possible to get around it. E.g. some gun you…

All of these platforms except perhaps Reddit are using LLMs (and other ML/AI) for censoring and automated anti-abuse.

Including the LLM platforms themselves.

Manual reporting is an adjunct/additional method, and goes into the training data set after whatever manual intervention occurs too.

Post reply on HN