Live data from Hacker News

AI behavior guardrails should be public

twitter.com

161–170 of 350 posts

Re: AI behavior guardrails should be public

#162

gemini seems to have problems generating white people and honestly this just opens the door for things that are even more racist [1], the harder you try the more you'll fail, just get over the DEI nonsense already 1. https://twitter.com/wagieeacc/status/1760371304425762940

I don't think the DEI stuff is nonsense, but SV is sensitive to this because most of their previous generation of models were horrifyingly racist if not teenage nazis, and so they turned the anti-racism knob up to 11 which made the models....racist but in a different way. Like depicting colonial settlers as native americans is extremely problematic in its own special way, but I also don't expect a statistical solver to grasp that context meaningfully.

Re: AI behavior guardrails should be public

#163
post #112

Harris and who I think was either Hughes or Stewart a podcast where they talked about how cringey and out of touch the elite are on the topic of race or wokeness in general. This faux pas on google's part couldn't be a better illustration of this. A bunch of wealthy rich tech geeks programming an AI to show racial diversity in what were/are unambiguously not diverse settings. They're just so painfully divorced from r…

I've found that anyone who uses the term "wokeness" seriously is likely arguing from a place of bad faith. It's origins are as a derogatory term, which people wanting to speak seriously on the topic should know.

Its origin was as a proud self-assigned term. It became derogatory entirely due to the behavior of said people. People wanting to speak seriously on the topic should avoid tone-policing and arguing about labels rather than the object referenced, despite knowing full well what is meant (otherwise, one wouldn't take offence)

Re: AI behavior guardrails should be public

#164
post #12

I strongly suspect Google tried really, really hard here to overcome the criticism is got with previous image recognition models saying that black people looked like gorillas. I am not really sure what I would want out of an image generation system, but I think Google's system probably went too far in trying to incorporate diversity in image generation.

Surely there is a middle ground. "Generate a scene of a group of friends enjoying lunch in the park." -> Totally expect racial and gender diversity in the output. "Generate a scene of 17th century kings of Scotland playing golf." -> The result should not be a bunch of black men and Asian women dressed up as Scottish kings, it should be a bunch of white guys.

There's no "middle" in the field of decompressing a short phrase into a visual scene (or program or book or whatever). There are countless private, implicit assumptions that users take for granted yet expect to see in the output, and vendors currently fear that their brand will be on the hook for the AI making a bad bet about those assumptions.

So for your first example, you totally expect racial and gender diversity in the output because you're assuming a realistic, contemporary, cosmopolitan, bourgeoisie setting -- either because you live in one or because you anticipate that the provider will default to one. The food will probably look Western, the friends will probably be young adults that look to have professional or service jobs wearing generic contemporary commercial fashion, the flora in in the park will be broadly northern climate, etc.

Most people around the world don't live in an environment anything like that, so nominal accuracy can't be what you're looking for. What you want, but don't say, is a scene that feels familiar to you and matches what you see as the de facto cultural ideal of contemporary Western society.

And conveniently, because a lot of the training data is already biased towards that society and the AI vendors know that the people who live in that society will be their most loyal customers and most dangerous critics right now, it's natural for them to put a thumb on the scale (through training, hidden prompts, etc) that gets the model to assume an innocuous Western-media-palatable middle ground -- so it delivers the racially and gender diverse middle class picnic in a generic US city park.

But then in your second example, you're implicitly asking for something historically accurate without actually saying that accuracy is what's become important for you in this new prompt. So the same thumb that biased your first prompt towards a globally-rare-but-customer-palatable contemporary, cosmopolitan, Western culture suddenly makes your new prompt produce something surreal and absurd.

There's no "middle" there because the problem is really in the unstated assumptions that we all carry into how we use these tools. It's more effective for them to make the default output Western-media-palatable and historical or cultural accuracy the exception that needs more explicit prompting.

If they're lucky, they may keep grinding on new training techniques and prompts that get more assumptions "right" by the people that matter to their success while still being inoffensive, but it's no simple "surely a middle ground" problem.

Re: AI behavior guardrails should be public

#165
post #66

Earlier quoted context omitted.

I remember checking like a year ago and they still had the word "gorilla" blacklisted (i.e. it never returns anything even if you have gorilla images).

Gotta love such a high quality fix. When your upper high tech, state of the art algorithm learns racist patterns just blocklist the word and move on. Don't worry about why it learned such patterns in the first place.

Humans do look like gorillas. We're related. It's natural that an imperfect program that deals with images will will mistake the two.

Humans, unfortunately, are offended if you imply they look like gorillas.

What's a good fix? Human sensitivity is arbitrary, so the fix is going to tend to be arbitrary too.

Re: AI behavior guardrails should be public

#166

Earlier quoted context omitted.

Surely there is a middle ground. "Generate a scene of a group of friends enjoying lunch in the park." -> Totally expect racial and gender diversity in the output. "Generate a scene of 17th century kings of Scotland playing golf." -> The result should not be a bunch of black men and Asian women dressed up as Scottish kings, it should be a bunch of white guys.

You can see how this gets challenging, though, right? If you train your model to prioritize real photos (as they're often more accurate representations than artistic ones), you might wind up with Denzel Washington as the archetype; https://en.wikipedia.org/wiki/The_Tragedy_of_Macbeth_(2021_f... . There's a vast gap between human understanding and what LLMs "understand".

I mean now you d to train AI to recognise the bias in the training data.

Re: AI behavior guardrails should be public

#167
post #83

Bing also generates political propaganda (guess of what side) if you ask it to generate images with the prompt "person holding a sign that says" without any further content. https://twitter.com/knn20000/status/1712562424845599045 https://twitter.com/ramonenomar/status/1722736169463750685 https://www.reddit.com/r/dalle2/comments/1ao1avd/why_did_thi... https://www.reddit.com/r/dalle2/comments/1ao1avd/why_did_thi...

As the images in your Reddit threads hilariously point out, you really shouldn’t believe everything you see on the internet, especially when it comes to AI generated content. Here is another example: https://www.thehour.com/entertainment/article/george-carlin-...

You should try yourself. The bing image generator is open and free. I tried the same prompts, and it is reproduceable. (Requires a few retries, though)

Re: AI behavior guardrails should be public

#168

I've never been involved with implementing large-scale moderation or content controls, but it seems pretty standard that underlying automated rules aren't generally public, and I've always assumed this is because there's a kind of necessary "security through obscurity" aspect to them. E.g., publish a word blocklist and people can easily find how to express problematic things using words that aren't on the list. Thing…

[flagged]

Re: AI behavior guardrails should be public

#169

Earlier quoted context omitted.

Gotta love such a high quality fix. When your upper high tech, state of the art algorithm learns racist patterns just blocklist the word and move on. Don't worry about why it learned such patterns in the first place.

Humans do look like gorillas. We're related. It's natural that an imperfect program that deals with images will will mistake the two. Humans, unfortunately, are offended if you imply they look like gorillas. What's a good fix? Human sensitivity is arbitrary, so the fix is going to tend to be arbitrary too.

A good fix would, in my opinion, understanding how the algorithm is actually categorizing and why it miss-recognized gorillas and humans.

If the algorithm doesn't work well they have problems to solve.

Re: AI behavior guardrails should be public

#170

I've never been involved with implementing large-scale moderation or content controls, but it seems pretty standard that underlying automated rules aren't generally public, and I've always assumed this is because there's a kind of necessary "security through obscurity" aspect to them. E.g., publish a word blocklist and people can easily find how to express problematic things using words that aren't on the list. Thing…

[flagged]

This feels more like a personal attack than a response to the argument made.
Post reply on HN