Live data from Hacker News

AI behavior guardrails should be public

twitter.com

41–50 of 350 posts

Re: AI behavior guardrails should be public

#41
post #7

I would also love to see more transparency around AI behavior guardrails, but I don't expect that will happen anytime soon. Transparency would make it much easier to circumvent guardrails.

Why is it an issue that you can circumvent the guardrails? I never understood that. The guard rails are there so that innocent people doesn't get bad responses with porn or racism, a user looking for porn or racism getting that doesn't seem to be a big deal.

>The guard rails are there so that innocent people doesn't get bad responses

The guardrails are also there so bad actors can't use the most powerful tools to generate deepfakes, disinformation videos and racist manifestos.

That Pandora's box will be open soon when local models run on cell phones and workstations with current datacenter-scale performance. I'm the meantime, they're holding back the tsunami of evil shit that will occur when AI goes uncontrolled.

Re: AI behavior guardrails should be public

#43
post #3

Curious to see if this thread gets flagged and shut down like the others. Shame, too, since I feel like all the Gemini stuff that’s gone down today is so important to talk about when we consider AI safety. This has convinced me more and more that the only possible way forward that’s not a dystopian hellscape is total freedom of all AI for anyone to do with as they wish. Anything else is forcing values on other people…

"The only way to deal with some people making crazy rules is to have no rules at all" --libertarians

"Oh my god I'm being eaten by a fucking bear" --also libertarians

Re: AI behavior guardrails should be public

#44

Earlier quoted context omitted.

Yeah, although it is weird that it doesn’t insert white people into results like this by accident? https://x.com/imao_/status/1760159905682509927?s=46 I’ve also seen numerous examples where it outright refuses to draw white people but will draw black people: https://x.com/iamyesyouareno/status/1760350903511449717?s=46 That doesn’t explainable by system prompt

Think about the training data. If the word "Zulu" appears in a label, it will be a non-White person 100% of the time. If the word "English" appears in a label, it will be a non-White person 10%+ of the time. Only 75% of modern England is White and most images in the training data were taken in modern times. Image models do not have deep semantic understanding yet. It is an LLM calling an Image model API. So "English"…

It inserts mostly colored people when you ask for Japanese as well, it isn't just the dataset.

Re: AI behavior guardrails should be public

#45

Harris and who I think was either Hughes or Stewart a podcast where they talked about how cringey and out of touch the elite are on the topic of race or wokeness in general. This faux pas on google's part couldn't be a better illustration of this. A bunch of wealthy rich tech geeks programming an AI to show racial diversity in what were/are unambiguously not diverse settings. They're just so painfully divorced from r…

The behaviour seems perfectly reasonable to me. They are not in the business of reflecting reality, they are in the business of creating it. To me what you call wokeness seems like a pretty good improvement

Re: AI behavior guardrails should be public

#46
post #3

Curious to see if this thread gets flagged and shut down like the others. Shame, too, since I feel like all the Gemini stuff that’s gone down today is so important to talk about when we consider AI safety. This has convinced me more and more that the only possible way forward that’s not a dystopian hellscape is total freedom of all AI for anyone to do with as they wish. Anything else is forcing values on other people…

I'm convinced this happens because of technical alignment challenges rather than a desire to present 1800s English Kings as non-white. > Use all possible different descents with equal probability. Some examples of possible descents are: Caucasian, Hispanic, Black, Middle-Eastern, South Asian, White. They should all have equal probability. This is OpenAI's system prompt. There is nothing nefarious here, they're asking…

> As these systems get better, they'll figure out that "1800s English" should mean "White with > 99.9% probability".

The thing is, they already could do that, if they weren't prompt engineered to do something else. The cleaner solution would be to let people prompt engineer such details themselves, instead of letting a US American company's idiosyncratic conception of "diversity" do the job. Japanese people would probably simply request "a group of Japanese people" instead of letting the hidden prompt modify "a group of people", where the US company unfortunately forgot to mention "East Asian" in their prompt apart from "South Asian".

Re: AI behavior guardrails should be public

#47

gemini seems to have problems generating white people and honestly this just opens the door for things that are even more racist [1], the harder you try the more you'll fail, just get over the DEI nonsense already 1. https://twitter.com/wagieeacc/status/1760371304425762940

Is there any evidence that this is a consequence of DEI rather than a deeper technical issue?

Re: AI behavior guardrails should be public

#48
post #44

Earlier quoted context omitted.

Think about the training data. If the word "Zulu" appears in a label, it will be a non-White person 100% of the time. If the word "English" appears in a label, it will be a non-White person 10%+ of the time. Only 75% of modern England is White and most images in the training data were taken in modern times. Image models do not have deep semantic understanding yet. It is an LLM calling an Image model API. So "English"…

It inserts mostly colored people when you ask for Japanese as well, it isn't just the dataset.

Yes it's a combination of blunt instrument system prompting + training data + cherry picking

Re: AI behavior guardrails should be public

#49
post #7

Earlier quoted context omitted.

Why is it an issue that you can circumvent the guardrails? I never understood that. The guard rails are there so that innocent people doesn't get bad responses with porn or racism, a user looking for porn or racism getting that doesn't seem to be a big deal.

The problem is bad actors who think porn or racism are intolerable in any form, who will publish mountains of articles condemning your chatbot for producing such things, even if they had to go out of their way to break the guardrails to make it do so. They will create boycotts against you, they will lobby government to make your life harder, they will petition payment processors and cloud service providers to not wor…

[deleted]

Re: AI behavior guardrails should be public

#50
post #7

Earlier quoted context omitted.

Why is it an issue that you can circumvent the guardrails? I never understood that. The guard rails are there so that innocent people doesn't get bad responses with porn or racism, a user looking for porn or racism getting that doesn't seem to be a big deal.

>The guard rails are there so that innocent people doesn't get bad responses The guardrails are also there so bad actors can't use the most powerful tools to generate deepfakes, disinformation videos and racist manifestos. That Pandora's box will be open soon when local models run on cell phones and workstations with current datacenter-scale performance. I'm the meantime, they're holding back the tsunami of evil shit…

No legal or financial strategist at OpenAI or Google is going to be worried about buying a couple months or years of fewer deepfakes out in the world as a whole.

Their concern is liability and brand. With the opportunity to stake out territory in an extremely promising new market, they don't want their brand associated with anything awkward to defend right now.

There may be a few idealist stewards who have the (debatable) anxieties you do and are advocating as you say, but they'd still need to be getting sign off from the more coldly strategic $$$$$ people.

Post reply on HN