Live data from Hacker News

AI behavior guardrails should be public

twitter.com

311–320 of 350 posts

Re: AI behavior guardrails should be public

#311
post #31
post #2

They know that people would be up in arms if it generated white men when you asked for black women so they went the safe route, but we need to show that the current result shouldn't be acceptable either.

See the prompt from yesterday's article on HN about the ChatGPT outage.[1] For example, all of a given occupation should not be the same gender or race. ... Use all possible different descents with equal probability. Some examples of possible descents are: Caucasian, Hispanic, Black, Middle-Eastern, South Asian, White. They should all have equal probability. Not the distribution that exists in the population. [1] htt…

Why do you assume this is the system prompt, and not a hallucination?

Re: AI behavior guardrails should be public

#312
post #225

Earlier quoted context omitted.

Its origin was as a proud self-assigned term. It became derogatory entirely due to the behavior of said people. People wanting to speak seriously on the topic should avoid tone-policing and arguing about labels rather than the object referenced, despite knowing full well what is meant (otherwise, one wouldn't take offence)

While the terms "woke," "stay woke," and similar are used to self describe by traditionally marginalized groups, the forms "wokeness" and "woke agenda" are predominately used outside these communities as a pejorative. https://en.wikipedia.org/wiki/Cultural_Marxism_conspiracy_th... https://www.inquirer.com/opinion/woke-bill-maher-olympics-re...

This is true, but you could say the same about "Tory", or many other labels for political groups that are used by supporters and opponents alike. It refers to a thing that some people have good and some people have poor opinions on, but the label is just a proxy, not a cause or carrier of opinion itself

Re: AI behavior guardrails should be public

#313

I've never been involved with implementing large-scale moderation or content controls, but it seems pretty standard that underlying automated rules aren't generally public, and I've always assumed this is because there's a kind of necessary "security through obscurity" aspect to them. E.g., publish a word blocklist and people can easily find how to express problematic things using words that aren't on the list. Thing…

Repeat after me: security by obscurity is weak. Clearly people can work out what some of the rules are, so why not just publish them. If you need to alter them when people figure out how to get around them, well, you already had to anyways.

Kerckoff's principle only applies to crypto. "Security by obscurity" is used in oodles of systems security applications and broader contexts involving human behavior.

Very very ordinary security best practices rely on obscurity. ALSR is a good example. It is defeated by a data exfiltration vulnerability but remains a useful thing to add to your binaries. Because outside of the crypto space security is an onion and layers add additional cost to attackers.

Re: AI behavior guardrails should be public

#314

Earlier quoted context omitted.

Surely there is a middle ground. "Generate a scene of a group of friends enjoying lunch in the park." -> Totally expect racial and gender diversity in the output. "Generate a scene of 17th century kings of Scotland playing golf." -> The result should not be a bunch of black men and Asian women dressed up as Scottish kings, it should be a bunch of white guys.

> Surely there is a middle ground. "Generate a scene of a group of friends enjoying lunch in the park." -> Totally expect racial and gender diversity in the output. Do we expect this because diverse groups are realistically most common or because we wish that they were? For example only some 10% of marriages are interracial, but commercials on TV would lead you to believe it’s 30% or higher. The goal for commercials…

19% of new marriages in 2019 (and likely to rise): https://en.m.wikipedia.org/wiki/Interracial_marriage_in_the_...

Plus, it’s still a recent change: Loving v Virginia (legalized interracial marriage across US) was decided in 1967.

Re: AI behavior guardrails should be public

#315
post #308

Earlier quoted context omitted.

Can you please explain how outright refusing to draw an image with from the prompt "white male scientist", and instead giving a lecture on how their race is irrelevant to their occupation, but then happily drawing the requested image when prompted for "black female scientist", is promoting inclusion and equality?

It is pretty clear to me. Reality has a bias, most scientist in the world are white males. This IA is overtuned in the opossite direction to inspire kids who have not ever seen a person like them, not white, in those kind of jobs.

Saying most scientists in the world are white males seems like a very Anglo-centric perspective, at least based on the numbers available from statista.com.

Re: AI behavior guardrails should be public

#316

Earlier quoted context omitted.

Why would this be flagged / shut down? Also, what Gemini stuff are you referring to?

This post reporting on the issue was https://news.ycombinator.com/item?id=39443459 Posts criticizing "DEI" measures (or even stating that they do exist) get flagged quite a lot

Wrong link? Nothing looks flagged

Re: AI behavior guardrails should be public

#317

Earlier quoted context omitted.

Repeat after me: security by obscurity is weak. Clearly people can work out what some of the rules are, so why not just publish them. If you need to alter them when people figure out how to get around them, well, you already had to anyways.

Kerckoff's principle only applies to crypto. "Security by obscurity" is used in oodles of systems security applications and broader contexts involving human behavior. Very very ordinary security best practices rely on obscurity. ALSR is a good example. It is defeated by a data exfiltration vulnerability but remains a useful thing to add to your binaries. Because outside of the crypto space security is an onion and la…

ASLR is not security by obscurity. The addresses into which all the various things are mapped are secret in the same sort of way as cryptographic secret and private keys are secret, but the mechanism is not secret.

Re: AI behavior guardrails should be public

#318
post #303

Earlier quoted context omitted.

Watch your opinion on this get silenced in subtle ways. From gaslighting to thread nerfing to vote locking.... Ask why anyone would engage in those behaviours vs the merit of the arguments and the voice of the people. The strings are revealing themselves so incredibly fast. edit: my first flagged! silence is deafening ^_^. This is achieved by nerfing the thread from public view, then allow the truly caustic to alter…

Could you please stop posting unsubstantive comments and flamebait and otherwise breaking the site guidelines? You've unfortunately been doing it repeatedly. If you wouldn't mind reviewing https://news.ycombinator.com/newsguidelines.html and taking the intended spirit of the site more to heart, we'd be grateful.

I’ll be more subtle in my agenda to match the spirit of the site

Re: AI behavior guardrails should be public

#319
post #297

Earlier quoted context omitted.

> They may get you arrested for trying to get into your own home The AI algorithms that ban people from platforms like Google and GitHub do that, which I explicitly called out as needing more oversight. That is different from algorithms which just prevent you from doing something on a platform like using n-bombs in your username, or the LLM guardrails that just give you mangled answers or tell you that they can't do…

Friend, it was your analogy

I was intending to make a point about security through obscurity, not to make an analogy with those systems.

Re: AI behavior guardrails should be public

#320
post #12

I strongly suspect Google tried really, really hard here to overcome the criticism is got with previous image recognition models saying that black people looked like gorillas. I am not really sure what I would want out of an image generation system, but I think Google's system probably went too far in trying to incorporate diversity in image generation.

They have now added a strong bias for generating black people now. Some have prompted to generate a picture of a German WW2 soldier, and now there are many pictures of black people floating around in NAZI uniforms. I think their strategy to "enhance" outcomes is very misdirected. The most widely used base models to really fine tune models are those that are not censored and I think you have to construct a problem to…

> ...when users are able to adapt models to their liking.

Therein lies the rub, as it were, because the large providers of AI models are working hard to ensure legislation that wouldn't allow people access to uncensored models in the name of "safety." And "safety" in this case includes the notion that models may not push the "correct" world-view enough.

Post reply on HN