Live data from Hacker News

AI behavior guardrails should be public

twitter.com

11–20 of 350 posts

Re: AI behavior guardrails should be public

#11
post #2

They know that people would be up in arms if it generated white men when you asked for black women so they went the safe route, but we need to show that the current result shouldn't be acceptable either.

the models are perfectly capable of generating exactly what they're told to.

instead, they covertly modify the prompts to make every request imaginable represent the human menagerie we're supposed to live in.

the results are hilarious. https://i.4cdn.org/g/1708514880730978.png

Re: AI behavior guardrails should be public

#12
I strongly suspect Google tried really, really hard here to overcome the criticism is got with previous image recognition models saying that black people looked like gorillas. I am not really sure what I would want out of an image generation system, but I think Google's system probably went too far in trying to incorporate diversity in image generation.

Re: AI behavior guardrails should be public

#13
post #6

I would also love to see more transparency around AI behavior guardrails, but I don't expect that will happen anytime soon. Transparency would make it much easier to circumvent guardrails.

Transparency may also subject these companies to litigation from groups that feel they are misrepresented in whatever way in the model.

This makes me wonder, how much lawyering is involved in the development of these tools?

Re: AI behavior guardrails should be public

#14
post #3

Curious to see if this thread gets flagged and shut down like the others. Shame, too, since I feel like all the Gemini stuff that’s gone down today is so important to talk about when we consider AI safety. This has convinced me more and more that the only possible way forward that’s not a dystopian hellscape is total freedom of all AI for anyone to do with as they wish. Anything else is forcing values on other people…

Why would this be flagged / shut down? Also, what Gemini stuff are you referring to?

[flagged]

Re: AI behavior guardrails should be public

#15
post #3

Curious to see if this thread gets flagged and shut down like the others. Shame, too, since I feel like all the Gemini stuff that’s gone down today is so important to talk about when we consider AI safety. This has convinced me more and more that the only possible way forward that’s not a dystopian hellscape is total freedom of all AI for anyone to do with as they wish. Anything else is forcing values on other people…

I'm convinced this happens because of technical alignment challenges rather than a desire to present 1800s English Kings as non-white.

> Use all possible different descents with equal probability. Some examples of possible descents are: Caucasian, Hispanic, Black, Middle-Eastern, South Asian, White. They should all have equal probability.

This is OpenAI's system prompt. There is nothing nefarious here, they're asking White to be chosen with high probability (Caucasian + White / 6 = 1/3) which is significantly more than how they're distributed in the general population.

The data these LLMs were trained on vastly over-represents wealthy countries who connected to the internet a decade earlier. If you don't explicitly put something in the system prompt, any time you ask for a "person" it will probably be Male and White, despite Male and White only being about 5-10% of the world's population. I would say that's even more dystopian. That the biases in the training distribution get automatically built-in and cemented forever unless we take active countermeasures.

As these systems get better, they'll figure out that "1800s English" should mean "White with > 99.9% probability". But as of February 2024, the hacky way we are doing system prompting is not there yet.

Re: AI behavior guardrails should be public

#16
post #6

Earlier quoted context omitted.

Transparency may also subject these companies to litigation from groups that feel they are misrepresented in whatever way in the model.

This makes me wonder, how much lawyering is involved in the development of these tools?

I often wonder if corporate lawyers just tell tech founders whatever they want to hear.

At a previous healthcare startup our founder asked us to build some really dodgy stuff with healthcare data. He assured us that it "cleared legal", but from everything I could tell it was in direct violation of the local healthcare info privacy acts.

I chose to find a new job at the time.

Re: AI behavior guardrails should be public

#18
post #7

Earlier quoted context omitted.

Why is it an issue that you can circumvent the guardrails? I never understood that. The guard rails are there so that innocent people doesn't get bad responses with porn or racism, a user looking for porn or racism getting that doesn't seem to be a big deal.

The problem is bad actors who think porn or racism are intolerable in any form, who will publish mountains of articles condemning your chatbot for producing such things, even if they had to go out of their way to break the guardrails to make it do so. They will create boycotts against you, they will lobby government to make your life harder, they will petition payment processors and cloud service providers to not wor…

But I can find porn and racism using Google search right now, how is that different? You have to disable their filters, but you can find it. Why is there no such thing for the google generation bots, I don't see why it would be so much worse here?

Re: AI behavior guardrails should be public

#20
post #9

Earlier quoted context omitted.

Why would this be flagged / shut down? Also, what Gemini stuff are you referring to?

Carmack’s tweet is about what’s going around Twitter today regarding the implicit biases Gemini (Google’s chatbot) has when drawing images. Will refuse to draw white people (and perhaps more strongly so, refuses to draw white men?) even in prompts where appropriate, like “Draw me a Pope” where Gemini drew an Indian woman and a Black man - here’s the thread: https://x.com/imao_/status/1760093853430710557?s=46 Maybe in…

EDIT: Nevermind.
Post reply on HN