Live data from Hacker News

AI behavior guardrails should be public

twitter.com

31–40 of 350 posts

Re: AI behavior guardrails should be public

#31
post #2

They know that people would be up in arms if it generated white men when you asked for black women so they went the safe route, but we need to show that the current result shouldn't be acceptable either.

See the prompt from yesterday's article on HN about the ChatGPT outage.[1]

For example, all of a given occupation should not be the same gender or race. ... Use all possible different descents with equal probability. Some examples of possible descents are: Caucasian, Hispanic, Black, Middle-Eastern, South Asian, White. They should all have equal probability.

Not the distribution that exists in the population.

[1] https://pastebin.com/vnxJ7kQk

Re: AI behavior guardrails should be public

#32
post #18

Earlier quoted context omitted.

The problem is bad actors who think porn or racism are intolerable in any form, who will publish mountains of articles condemning your chatbot for producing such things, even if they had to go out of their way to break the guardrails to make it do so. They will create boycotts against you, they will lobby government to make your life harder, they will petition payment processors and cloud service providers to not wor…

But I can find porn and racism using Google search right now, how is that different? You have to disable their filters, but you can find it. Why is there no such thing for the google generation bots, I don't see why it would be so much worse here?

I'm leaning towards 'there is a difference between being the one who enables access to x and being the one who created x' (albeit not a substantive one for the end user), but that leaves open the question of why that doesn't apply to, eg, social media platforms. Maybe people think of google search as closer to an ISP than a platform?

Re: AI behavior guardrails should be public

#34
Harris and who I think was either Hughes or Stewart a podcast where they talked about how cringey and out of touch the elite are on the topic of race or wokeness in general.

This faux pas on google's part couldn't be a better illustration of this. A bunch of wealthy rich tech geeks programming an AI to show racial diversity in what were/are unambiguously not diverse settings.

They're just so painfully divorced from reality that they are just acting as a multiplier in making the problem worse. People say that we on the left are driving around a clown car, and google is out their putting polka dots and squeaky horns on the hood.

Re: AI behavior guardrails should be public

#36
post #18

Earlier quoted context omitted.

The problem is bad actors who think porn or racism are intolerable in any form, who will publish mountains of articles condemning your chatbot for producing such things, even if they had to go out of their way to break the guardrails to make it do so. They will create boycotts against you, they will lobby government to make your life harder, they will petition payment processors and cloud service providers to not wor…

But I can find porn and racism using Google search right now, how is that different? You have to disable their filters, but you can find it. Why is there no such thing for the google generation bots, I don't see why it would be so much worse here?

> how is that different?

Because 'those' legal battles over search have already been fought and are established law across most countries.

When you throw in some new application now all that same stuff goes back to court and gets fought again. Section 230 is already legally contentious enough these days.

Re: AI behavior guardrails should be public

#37

Harris and who I think was either Hughes or Stewart a podcast where they talked about how cringey and out of touch the elite are on the topic of race or wokeness in general. This faux pas on google's part couldn't be a better illustration of this. A bunch of wealthy rich tech geeks programming an AI to show racial diversity in what were/are unambiguously not diverse settings. They're just so painfully divorced from r…

Watch your opinion on this get silenced in subtle ways. From gaslighting to thread nerfing to vote locking.... Ask why anyone would engage in those behaviours vs the merit of the arguments and the voice of the people.

The strings are revealing themselves so incredibly fast.

edit: my first flagged! silence is deafening ^_^. This is achieved by nerfing the thread from public view, then allow the truly caustic to alter the vote ratio in a way that makes opinion appear more balanced than it really is. Nice work, kleptomaniacs

Re: AI behavior guardrails should be public

#38

Earlier quoted context omitted.

I'm convinced this happens because of technical alignment challenges rather than a desire to present 1800s English Kings as non-white. > Use all possible different descents with equal probability. Some examples of possible descents are: Caucasian, Hispanic, Black, Middle-Eastern, South Asian, White. They should all have equal probability. This is OpenAI's system prompt. There is nothing nefarious here, they're asking…

Yeah, although it is weird that it doesn’t insert white people into results like this by accident? https://x.com/imao_/status/1760159905682509927?s=46 I’ve also seen numerous examples where it outright refuses to draw white people but will draw black people: https://x.com/iamyesyouareno/status/1760350903511449717?s=46 That doesn’t explainable by system prompt

Think about the training data.

If the word "Zulu" appears in a label, it will be a non-White person 100% of the time.

If the word "English" appears in a label, it will be a non-White person 10%+ of the time. Only 75% of modern England is White and most images in the training data were taken in modern times.

Image models do not have deep semantic understanding yet. It is an LLM calling an Image model API. So "English" + "Kings" are treated as separate conceptual things, then you get 5-10% of the results as non-White people as per its training data.

https://postimg.cc/0zR35sC1

Add to this massive amounts of cherry picking on "X", and you get this kind of bullshit culture war outrage.

I really would have expected technical people to be better than this.

Re: AI behavior guardrails should be public

#39
post #7

I would also love to see more transparency around AI behavior guardrails, but I don't expect that will happen anytime soon. Transparency would make it much easier to circumvent guardrails.

Why is it an issue that you can circumvent the guardrails? I never understood that. The guard rails are there so that innocent people doesn't get bad responses with porn or racism, a user looking for porn or racism getting that doesn't seem to be a big deal.

If you can get it on purpose, you can get it on accident. There's no perfect filter available so companies choose to cut more and stay on the safe side. It's not even just the overt cases - their systems are used by businesses and getting a bad response is a risk. Think of the recent incident with airline chatbot giving wrong answers. Now think of the cases where GPT gave racially biased answers in code as an example.

As a user who makes any business decision or does user communication including LLM, you really don't want to have a bad day because the LLM learned about some bias decided to merge it into your answer.

Re: AI behavior guardrails should be public

#40
gemini seems to have problems generating white people and honestly this just opens the door for things that are even more racist [1], the harder you try the more you'll fail, just get over the DEI nonsense already

1. https://twitter.com/wagieeacc/status/1760371304425762940

Post reply on HN