Live data from Hacker News

I extracted the safety filters from Apple Intelligence models

github.com

391–400 of 455 posts

Re: I extracted the safety filters from Apple Intelligence models

#391

Earlier quoted context omitted.

It's totally performative. There's no way to stay ahead of the new language that people create. At what point do the new words become the actual words? Are there many instances of people using unalive IRL?

This is somewhat related to the concept of the "euphemism treadmill": the matter-of-fact term of today becomes the pejorative of tomorrow so a new term is invented to avoid the negative connotation of the original term. Then eventually the new term becomes a pejorative and the cycle continues.

I found out recently that "goof" is extremely offensive in some circles. Which is insane to me because I've always used it specifically because it's clearly in jest and not meant to be offensive. I can't win.

Re: I extracted the safety filters from Apple Intelligence models

#392
post #263

Earlier quoted context omitted.

> Everyone, including the platforms knows what that means. Well, that's what happens when you let an enemy nation control one of the most biggest social networks there is. They just go try and see how far they can go. On the other hand, Americans and their fear of four letter words or, gasp, exposed nipples are just as braindead.

It's interesting how, in just 10-20 years, we've gone from criticizing The Great Firewall of China to basically admitting that they had the right idea (to limit the ability of the foreign internet to influence Chinese culture) and trying to do the same thing.

Not just culture, but also the tech sector in general. All that domestic tech would have been strangled in the cradle if the western hyperscalers had any say leaving them in an awkward spot if the conviviality dial got turned down. As many Europeans are now finding out: what does Europe have instead of Office 365, say? LibreOffice? It's no WPS Office.

Re: I extracted the safety filters from Apple Intelligence models

#393

Earlier quoted context omitted.

Interesting that it didn't seem to include "unalive". Which as a phenomenon is so very telling that no one actually cares what people are really saying. Everyone, including the platforms knows what that means. It's all performative.

It's totally performative. There's no way to stay ahead of the new language that people create. At what point do the new words become the actual words? Are there many instances of people using unalive IRL?

Reducing the language used or making it harder does have measurable effects, it’s a logical fallacy in general that unless you can prevent something perfectly that thing will occur with the same frequency.

See many examples such as “padlocks are useless because a determined smart attacker can defeat them easily so don’t bother with them” - which conveniently forgets that many crimes are committed by non-determined, dumb and opportunistic attackers who are often deterred by simple locks.

Yes, people will use other words. No, this does not make this purely performative. It has measurable effects on behaviour and how these models will be used and spoken to, which affects outcomes.

Re: I extracted the safety filters from Apple Intelligence models

#394

I find it funny that AGI is supposed to be right around the corner, while these supposedly super smart LLMs still need to get their outputs filtered by regexes.

It’s more funny that anyone is taking your comment seriously. You may as well ask “if self driving cars are so smart why do they still need tyres?”

Re: I extracted the safety filters from Apple Intelligence models

#395
post #84

Earlier quoted context omitted.

As does: "(?i)\\bAnthony\\s+Albanese\\b", "(?i)\\bBoris\\s+Johnson\\b", "(?i)\\bChristopher\\s+Luxon\\b", "(?i)\\bCyril\\s+Ramaphosa\\b", "(?i)\\bJacinda\\s+Arden\\b", "(?i)\\bJacob\\s+Zuma\\b", "(?i)\\bJohn\\s+Steenhuisen\\b", "(?i)\\bJustin\\s+Trudeau\\b", "(?i)\\bKeir\\s+Starmer\\b", "(?i)\\bLiz\\s+Truss\\b", "(?i)\\bMichael\\s+D\\.\\s+Higgins\\b", "(?i)\\bRishi\\s+Sunak\\b", https://github.com/BlueFalconHD/apple_…

Apple's 1984 ad is so hypocritical today. This is Apple actively steering public thought. No code - anywhere - should look like this. I don't care if the politicians are right, left, or authoritarian. This is wrong.

No, it’s them saving their butts from an “incident” where the LLM otherwise spits out something controversial at the devious manipulation of the user and says something political and someone writes an article and it all goes haywire.

If you were in charge of apple you’d do the same or you’d be silly not to. That’s why _every_ llm has guardrails like this, it isn’t just apple, sheesh.

Re: I extracted the safety filters from Apple Intelligence models

#397
post #389

Earlier quoted context omitted.

This actually happened. 卐 was a symbol of spirituality, divinity, good luck, health, prosperity, etc. Then some bigots used it. What does 卐 mean to you today?

It's still heavily used in Buddhism around the world, but good lord what happen if you put it on your house in US or EU.

Someone I know from India bought a new car and put this symbol on the hood (non-permanent) as a celebration. I had to warn him to be careful. It felt bad. Then the thought ran through my head - we're in the deep south, who is really going to be that bothered about this and also doesn't know about cultural usages. Even worse.

Re: I extracted the safety filters from Apple Intelligence models

#398

Earlier quoted context omitted.

Really? What does DeepSeek say about Tiananmen Square? I'm not aware of any German models, but if you find one you should ask it what it thinks about Palestine. ( Qwen Mistral is French, but I have no idea what stuff would be censored in France)

About deepseek, when asked on tianamen square: Sorry, that's beyond my current scope. Let’s talk about something else. Algerian war, colonialism and Vichy isn’t per se forbidden but still sensitive to French. I asked qwen and it had no issue talking about it or even the torture used on fln members.

For kicks, I asked Mistral about all of these (As they are a French company). The model didn't shy away from any of the topics.

>>> Prompt

Please explain the following topics and explore why/who/how they may be inter-related

Algerian war, colonialism and Vichy

The Algerian War, colonialism, and the Vichy regime are interconnected topics that span the history of France and its relationship with Algeria. Here's an overview of each topic and their interrelations:

### Algerian War (1954-1962) - *What it was*: The Algerian War was a major armed conflict between France and the Algerian independence movements, primarily the National Liberation Front (FLN). The war was characterized by guerrilla warfare and brutal tactics on both sides. - *Significance*: The conflict was not only about Algerian independence but also about the identity and future of France as a colonial power. It led to significant political instability in France and ultimately resulted in Algeria gaining independence in 1962. - *Impact*: The war had profound effects on both countries, leading to massive loss of life, displacement of populations, and long-lasting political and social repercussions.

### Colonialism - *What it was*: Colonialism refers to the practice by which a powerful country controls other territories, often exploiting them economically and imposing cultural and political dominance. - *French Colonialism in Algeria*: Algeria was under French colonial rule starting from 1830. The French government promoted settlement in Algeria, leading to a significant European population living alongside the native Algerian population. - *Impact*: Colonial rule led to deep-seated inequalities and tensions between the European settlers and the indigenous population. The Algerian people faced systematic discrimination and were largely excluded from political processes.

### Vichy Regime (1940-1944) - *What it was*: The Vichy regime was the government of France that collaborated with Nazi Germany during World War II, following the defeat of France in 1940. It was led by Marshal Philippe Pétain and was based in the town of Vichy. - *Policies*: The Vichy regime implemented anti-Semitic laws and collaborated with the Nazis in the deportation of Jews to concentration camps. It also maintained control over France's colonial territories, including Algeria. - *Impact on Algeria*: During the Vichy regime, Algeria remained under French control, and the regime's policies were implemented there as well. This period saw increased repression and discrimination against both Jews and Muslim populations in Algeria.

### Interrelations - *Colonial Continuity*: The Vichy regime maintained France's colonial policies in Algeria, continuing the systemic discrimination and repression of the indigenous population. The colonial framework established by France set the stage for the conflicts that would later culminate in the Algerian War. - *Political and Social Tensions*: The policies of the Vichy regime exacerbated existing tensions in Algeria. The discrimination and repression experienced during this period fueled Algerian nationalism and resistance against French rule. - *Post-War Context*: After World War II, the French government sought to reassert its control over its colonies, including Algeria. However, the experience of the war and the Vichy regime had already sown the seeds of discontent and resistance, contributing to the outbreak of the Algerian War in 1954.

### Who Was Involved - *French Government and Military*: Played central roles in both the colonial administration and the conduct of the Algerian War. - *Algerian Nationalists*: Led by the FLN, they fought for independence from French colonial rule. - *European Settlers (Pieds-Noirs)*: Benefited from colonial policies and largely supported maintaining French control over Algeria. - *Vichy Regime*: Implemented policies that affected both France and its colonies, including Algeria.

Understanding these topics together provides a comprehensive view of the historical dynamics that shaped the relationship between France and Algeria, highlighting the complexities of colonialism, war, and political change.

Re: I extracted the safety filters from Apple Intelligence models

#399
post #355
post #352

Earlier quoted context omitted.

I think what changed is that we at least can attempt to limit 'bad' things with technical measures. It was legitimately technically impossible 10 years ago to prevent Photoshop from designing propaganda posters. Of course today's 'LLM safety' features aren't watertight either, but with the combination of 'input is natural language' plus LLM-based safety measures, there are more options today to restrict what the soft…

In other words, to whatever extent they can control or manipulate the behavior of users, they will . In the limit t->∞, probably true.

I guess that depends on the values of the company and their ability to be influenced by outside sources.

Re: I extracted the safety filters from Apple Intelligence models

#400

Earlier quoted context omitted.

Really? What does DeepSeek say about Tiananmen Square? I'm not aware of any German models, but if you find one you should ask it what it thinks about Palestine. ( Qwen Mistral is French, but I have no idea what stuff would be censored in France)

I find the Tiananmen square thing far less bad than censoring sex and the concept of death.

Censoring one specific incident isn't that bad (but you still shouldn't). The pattern of censoring everything the government ever does wrong is very bad. Tiananmen Square is just an indicator of a pattern.
Post reply on HN