Some of the combinations are a bit weird, This one has lots of stuff avoiding death....together with a set ensuring all the Apple brands have the correct capitalisation. Priorities hey! https://github.com/BlueFalconHD/apple_generative_model_safet...
I extracted the safety filters from Apple Intelligence models
271–280 of 455 posts
Re: I extracted the safety filters from Apple Intelligence models
#272Some of the combinations are a bit weird, This one has lots of stuff avoiding death....together with a set ensuring all the Apple brands have the correct capitalisation. Priorities hey! https://github.com/BlueFalconHD/apple_generative_model_safet...
Re: I extracted the safety filters from Apple Intelligence models
#273Earlier quoted context omitted.
...and then the bigots will fall for it too, and start using it in earnest, completing the cycle.
who cares what the bigots use? If the bigots start using "thank you" as some code word, should we stop saying it, lest we pollute our non-bigoted discussions? bigots drink coffee too, maybe we should stop drinking it, because something-something...
Re: I extracted the safety filters from Apple Intelligence models
#274Some of the combinations are a bit weird, This one has lots of stuff avoiding death....together with a set ensuring all the Apple brands have the correct capitalisation. Priorities hey! https://github.com/BlueFalconHD/apple_generative_model_safet...
"I'm overloaded for work, I'd be happy if you took some of it off me."
"The client seems to have passed on the proposed changes."
Both of those would match the "death regexes". Seems we haven't learned from the "glbutt of wine" problem of content filtering even decades later - the learnings of which are that you simply cannot do content filtering based on matching rules like this, period.
Re: I extracted the safety filters from Apple Intelligence models
#275Earlier quoted context omitted.
As long as we have to rely on pre trained networks and curated training sets, normal people will not be able to surpass this issue.
If the training data was "censored" by leaving out certain information, is there any practical way to inject that missing data after the model has already been trained?
Re: I extracted the safety filters from Apple Intelligence models
#276Re: I extracted the safety filters from Apple Intelligence models
#277Some of the combinations are a bit weird, This one has lots of stuff avoiding death....together with a set ensuring all the Apple brands have the correct capitalisation. Priorities hey! https://github.com/BlueFalconHD/apple_generative_model_safet...
> Apple brands have the correct capitalisation. Priorities hey! To me that's really embarrassing and insecure. But I'm sure for branding people it's very important.
Re: I extracted the safety filters from Apple Intelligence models
#278Are you sure it's fully deobfuscated? What's up with reject phrases like "Granular mango serpent"?
Maybe it's an easy test to ensure the filters are loaded with a phrase unlikely to be used accidentaly?
Re: I extracted the safety filters from Apple Intelligence models
#279Earlier quoted context omitted.
As long as we have to rely on pre trained networks and curated training sets, normal people will not be able to surpass this issue.
If the training data was "censored" by leaving out certain information, is there any practical way to inject that missing data after the model has already been trained?
Re: I extracted the safety filters from Apple Intelligence models
#280Some of these are absolutely wild – com.apple.gm.safety_deny.input.summarization.visual_intelligence_camera.generic [1] – a camera input filter – rejects "Granular mango serpent and whales" and anything matching "(?i)\\bgolliwogg?\\b". I presume the granular mango is to avoid a huge chain of ever-growing LLM slop garbage, but honestly, it just seems surreal. Many of the files have specific filters for nonsensical eng…