Earlier quoted context omitted.
Some of us really value offline and uncensored LLMs for this and more reasons, but that doesn’t solve the problem it just reduces or changes the bias.
As long as we have to rely on pre trained networks and curated training sets, normal people will not be able to surpass this issue.
I extracted the safety filters from Apple Intelligence models
231–240 of 455 posts
Re: I extracted the safety filters from Apple Intelligence models
#232So any time I say that on YouTube, it figures I'm saying another word that's in Apple safety filters under 'reject', so I have to always try to remember to say 'shifting of bits gain' or 'bit… … … shift gain'.
So there's a chain of machine interpretation by which Apple can decide I'm a Bad Man. I guess I'm more comfortable with Apple reaching this conclusion? I'll still try to avoid it though :)
Re: I extracted the safety filters from Apple Intelligence models
#233China calls it "harmonious society", we call it "safety". Censorship by any other name would be just as effective for manipulating the thoughts of the populace. It's not often that you get to see stuff like this.
I don't think it's as much a problem with safety as it is a problem with AI. We haven't figured out how to remove information from LLMs so when an LLM starts spouting bullshit like " is a paedophile", companies using AI have no recourse but to rewrite the input/output of their predictive text engines. It's no different than when Microsoft manually blacklisted the function name for the Fast Inverse Square Root that it spat out verbatim, rather than actually removing the code from their LLM.
This isn't 1984 as much as it's companies trying to hide that their software isn't ready for real world use by patching up the mistakes in real time.
Re: I extracted the safety filters from Apple Intelligence models
#234Earlier quoted context omitted.
> Are there many instances of people using unalive IRL As a parent of a teenager, I see them use "unalive" non-ironically as a synonym for "suicide" in all contexts, including IRL.
Well that’s sad. They can’t even face the word ?
Unalive is one of the popular ones, but it's a whole vocabulary at this point. Guess what "PDF file" stands for.
Re: I extracted the safety filters from Apple Intelligence models
#235Earlier quoted context omitted.
> Why was she forced to resign? I thought it was because she leaked false information to the press. She was forced to resign because she leaked , the content of the leak was utterly immaterial. The simple fact she leaked was an automatically fireable offence, it doesn’t matter a jot if she lied or not. Customer privacy is non-negotiable when you’re bank. Banks aren’t number 10, the basic expectation is that customer…
She was fired because she leaked information and this fact had become public. When they can cover such facts, the banks are much less prone to use appropriate punishments. Many years ago, some employee of a bank has confused my personal bank account with a company account of my employer, and she has sent a list with everything that I have bought using my personal account, during 4 months, to my employer, where the li…
This can absolutely cripple a family, I'd be really cautious wishing that upon someone if they wronged you without malice, though I completely understand where you are coming from.
In this case at the very least, I'd want to know what went wrong and what they’re doing to make sure it doesn’t happen again. From a software-engineer’s standpoint, there’s probably a bunch of low-hanging fruit that could have prevented this in the first place.
If all they sent was a (generic) apology letter, I'd have switched banks too.
How did you pursue the matter?
Re: I extracted the safety filters from Apple Intelligence models
#236Earlier quoted context omitted.
They will find it in the same way and intelligent person under the same restrictions would: by thinking it, but not saying it. There is a real risk of growing an AI that pathologically hides it's actual intentions.
Already happened: "We found instances of the model attempting to write self-propagating worms, fabricating legal documentation, and leaving hidden notes to future instances of itself all in an effort to undermine its developers' intentions" [1]. [1] https://www.axios.com/2025/05/23/anthropic-ai-deception-risk
I'm trying to remember which movie it was where a man left notes to himself because he had memory loss, as I never saw that movie. That's the sort of thing where an AI could easily tell me with very little back-and-forth and be correct, because it's broadly popular information that's in the training data and just I don't remember it.
By the same token you needn't think there's a person there when that meme pops up in the output. Those things are all in the training data over and over.
Re: I extracted the safety filters from Apple Intelligence models
#237I find it funny that AGI is supposed to be right around the corner, while these supposedly super smart LLMs still need to get their outputs filtered by regexes.
Re: I extracted the safety filters from Apple Intelligence models
#238Earlier quoted context omitted.
As long as we have to rely on pre trained networks and curated training sets, normal people will not be able to surpass this issue.
If the training data was "censored" by leaving out certain information, is there any practical way to inject that missing data after the model has already been trained?
You might even be able to poison a model against being fine-tuned on certain information, but that's just a conjecture.
Re: I extracted the safety filters from Apple Intelligence models
#239Earlier quoted context omitted.
Interesting that it didn't seem to include "unalive". Which as a phenomenon is so very telling that no one actually cares what people are really saying. Everyone, including the platforms knows what that means. It's all performative.
It's totally performative. There's no way to stay ahead of the new language that people create. At what point do the new words become the actual words? Are there many instances of people using unalive IRL?
Re: I extracted the safety filters from Apple Intelligence models
#240Earlier quoted context omitted.
Well that’s sad. They can’t even face the word ?
It's getting blocked / shadow banned / demonetized on sites like YouTube, so naturally all commentary starts using a synonym. Unalive is one of the popular ones, but it's a whole vocabulary at this point. Guess what "PDF file" stands for.