Live data from Hacker News

AI behavior guardrails should be public

twitter.com

21–30 of 350 posts

Re: AI behavior guardrails should be public

#21
post #2

They know that people would be up in arms if it generated white men when you asked for black women so they went the safe route, but we need to show that the current result shouldn't be acceptable either.

the models are perfectly capable of generating exactly what they're told to. instead, they covertly modify the prompts to make every request imaginable represent the human menagerie we're supposed to live in. the results are hilarious. https://i.4cdn.org/g/1708514880730978.png

If you're gonna take an image from /g/ and post it, upload it somewhere else first - 4chan posts deliberately go away after the thread gets bumped off. A direct link is going to rot very quickly.

Re: AI behavior guardrails should be public

#22
post #3

Curious to see if this thread gets flagged and shut down like the others. Shame, too, since I feel like all the Gemini stuff that’s gone down today is so important to talk about when we consider AI safety. This has convinced me more and more that the only possible way forward that’s not a dystopian hellscape is total freedom of all AI for anyone to do with as they wish. Anything else is forcing values on other people…

I'm convinced this happens because of technical alignment challenges rather than a desire to present 1800s English Kings as non-white. > Use all possible different descents with equal probability. Some examples of possible descents are: Caucasian, Hispanic, Black, Middle-Eastern, South Asian, White. They should all have equal probability. This is OpenAI's system prompt. There is nothing nefarious here, they're asking…

Yeah, although it is weird that it doesn’t insert white people into results like this by accident? https://x.com/imao_/status/1760159905682509927?s=46

I’ve also seen numerous examples where it outright refuses to draw white people but will draw black people: https://x.com/iamyesyouareno/status/1760350903511449717?s=46

That doesn’t explainable by system prompt

Re: AI behavior guardrails should be public

#23
post #6

Earlier quoted context omitted.

Transparency may also subject these companies to litigation from groups that feel they are misrepresented in whatever way in the model.

This makes me wonder, how much lawyering is involved in the development of these tools?

I've had 'AI Attorneys' on Twitter unable to even debate the most basic of arguments. It is definitely a self fulfilling death spiral and no one wants to check reality.

Re: AI behavior guardrails should be public

#24
post #18

Earlier quoted context omitted.

The problem is bad actors who think porn or racism are intolerable in any form, who will publish mountains of articles condemning your chatbot for producing such things, even if they had to go out of their way to break the guardrails to make it do so. They will create boycotts against you, they will lobby government to make your life harder, they will petition payment processors and cloud service providers to not wor…

But I can find porn and racism using Google search right now, how is that different? You have to disable their filters, but you can find it. Why is there no such thing for the google generation bots, I don't see why it would be so much worse here?

I cannot explain why Google gets a pass, possibly just because they are well entrenched and not an easy target.

But AI models are new, they are vulnerable to criticism, and they are absolutely ripe for a group of "antis" to form around.

Re: AI behavior guardrails should be public

#25
post #9

Earlier quoted context omitted.

Carmack’s tweet is about what’s going around Twitter today regarding the implicit biases Gemini (Google’s chatbot) has when drawing images. Will refuse to draw white people (and perhaps more strongly so, refuses to draw white men?) even in prompts where appropriate, like “Draw me a Pope” where Gemini drew an Indian woman and a Black man - here’s the thread: https://x.com/imao_/status/1760093853430710557?s=46 Maybe in…

EDIT: Nevermind.

[flagged]

Re: AI behavior guardrails should be public

#27
post #9

Earlier quoted context omitted.

Carmack’s tweet is about what’s going around Twitter today regarding the implicit biases Gemini (Google’s chatbot) has when drawing images. Will refuse to draw white people (and perhaps more strongly so, refuses to draw white men?) even in prompts where appropriate, like “Draw me a Pope” where Gemini drew an Indian woman and a Black man - here’s the thread: https://x.com/imao_/status/1760093853430710557?s=46 Maybe in…

EDIT: Nevermind.

It’s quite non-deterministic and it’s been patched since the middle of the day, as per a Google director https://x.com/jackk/status/1760334258722250785?s=46

Fwiw, it seems to have gone deeper than outright historical replacement: https://x.com/iamyesyouareno/status/1760350903511449717?s=46

Re: AI behavior guardrails should be public

#29
post #7

I would also love to see more transparency around AI behavior guardrails, but I don't expect that will happen anytime soon. Transparency would make it much easier to circumvent guardrails.

Why is it an issue that you can circumvent the guardrails? I never understood that. The guard rails are there so that innocent people doesn't get bad responses with porn or racism, a user looking for porn or racism getting that doesn't seem to be a big deal.

racism victims being defined in 2024 by anyone but western/white people. being erased seems ok. can you bet than in 20 years the standard will not shift to mixed race people like me? then you will also call people complaining racist and put guardrails against them... this is where it is going

Re: AI behavior guardrails should be public

#30

Earlier quoted context omitted.

Why would this be flagged / shut down? Also, what Gemini stuff are you referring to?

> Why would this be flagged / shut down A lot of people believe (based on a fair amount of evidence) that public AI tools like ChatGPT are forced by the guardrails to follow a particular (left-wing) script. There's no absolute proof of that, though, because they're kept a closely-guarded secret. These discussions get shut down when people start presenting evidence of baked-in bias.

The rationalization for injecting bias rests on two core ideas:

A. It is claimed that all perspectives are 'inherently biased'. There is no objective truth. The bias the actor injects is just as valid as another.

B. It is claimed that some perspectives carry an inherent 'harmful bias'. It is the mission of the actor to protect the world from this harm. There is no open definition of what the harm is and how to measure it.

I don't see how we can build a stable democratic society based on these ideas. It is placing too much power in too few hands. He who wields the levers of power, gets to define what biases to underpin the very basis of the social perception of reality, including but not limited to rewriting history to fit his agenda. There are no checks and balances.

Arguably there were never checks and balances, other than market competition. The trouble is that information technology and globalization have produced a hyper-scale society, in which, by Pareto's law, the power is concentrated in the hands of very few, at the helm of a handful global scale behemoths.

Post reply on HN