Live data from Hacker News

AI behavior guardrails should be public

twitter.com

201–210 of 350 posts

Re: AI behavior guardrails should be public

#201
post #187

Earlier quoted context omitted.

It'll also add extra fingers to human hands. Presumably that's not because of DEI guardrails about polydactyly, right? The current state of the art in AI gets things wrong regularly .

This specific thing is a much more blatant class of error, and one that has been known to occur in several previous models because of DEI systems (e.g. in cases where prompts have been leaked), and has never been known to occur for any other reason. Yes, it's conceivable that Google's newer, beter-than-ever-before AI system somehow has a fundamental technical problem that coincidentally just happens to cause the same…

> has never been known to occur for any other reason

Of course it has. Again, these things regularly give humans extra fingers and arms. They don't even know what humans fundamentally look like.

On the flip side, humans are shitty at recognizing bias. This comment thread stems from someone complaining the AI only rarely generated white people, but that's statistically accurate. It feels biased to someone in a majority-white nation with majority-white friends and coworkers, but it fundamentally isn't.

I don't doubt that there are some attempts to get LLMs to go outside the "white westerner" bubble in training sets and prompts. I suspect the extent of it is also deeply exaggerated by those who like to throw around woke-this and woke-that as derogatories.

Re: AI behavior guardrails should be public

#202
post #73

Earlier quoted context omitted.

> Where do we go from here? opensource models and training sets. So basically the "secret sauce" minus the hardware. I don't see it happening voluntarily.

I see it as not unlikely that there'll be a campaign to sigmatize, if not outright ban, open source models on the grounds of "safety". I'm quite surprised at how relatively unimpeded the distribution of image generation models has been, so far

This is already happening, actually, although the focus so far has been on the possibility of their use for CSAM:

https://www.theguardian.com/society/2023/sep/12/paedophiles-...

Re: AI behavior guardrails should be public

#203

Earlier quoted context omitted.

Surely there is a middle ground. "Generate a scene of a group of friends enjoying lunch in the park." -> Totally expect racial and gender diversity in the output. "Generate a scene of 17th century kings of Scotland playing golf." -> The result should not be a bunch of black men and Asian women dressed up as Scottish kings, it should be a bunch of white guys.

You can see how this gets challenging, though, right? If you train your model to prioritize real photos (as they're often more accurate representations than artistic ones), you might wind up with Denzel Washington as the archetype; https://en.wikipedia.org/wiki/The_Tragedy_of_Macbeth_(2021_f... . There's a vast gap between human understanding and what LLMs "understand".

> If you train your model to prioritize real photos

I thought that was the big bugbear about disinformation and false news, but now we have to censor reality to combat "bias"

Re: AI behavior guardrails should be public

#204

Earlier quoted context omitted.

There is no need to implement large scale censorship and moderation in this case. Where is the security concern? That I can generate images of white people in various situations for my five minutes of entertainment? The whole premise of your argument doesn't make sense. I'm talking to a computer, nobody gets hurt. It's like censoring what I write in my notes app vs. what I write on someone's Facebook wall. In one cas…

When these companies say there are "security concerns" they mean for them, not you! And they mean the security of their profits. So anything that can cause them legal liability or cause them brand degradation is a "security concern".

It's definitely this at this stage. But by not having any discourse we'll end up normalizing it even before establishing consensus on what's appropriate to expect from human/AI interaction and how much of a problem is the actual model, opposed to a user. Not being able to generate innocent content is ridiculous. Probably they're overshooting and learning how to draw the stricter lines right now, but if you don't argue, you'll allow to boil this frog into a Google "Search" again.

Re: AI behavior guardrails should be public

#205

I've never been involved with implementing large-scale moderation or content controls, but it seems pretty standard that underlying automated rules aren't generally public, and I've always assumed this is because there's a kind of necessary "security through obscurity" aspect to them. E.g., publish a word blocklist and people can easily find how to express problematic things using words that aren't on the list. Thing…

My dear fellow, some believe the ends justify the means and play games. Read history, have some decency.

The danger of being captured by such people far outweighs any other "problematic things".

First and foremost any system must defend against that. You love guardrails so much - put them on the self annointed guard railers.

Otherwise, if You Want a Picture of the Future, Imagine a Boot Stamping on a Human Face – for Ever.

Re: AI behavior guardrails should be public

#206
post #88

Earlier quoted context omitted.

Sure, but this one is from Google adding a tag to make every image of people diverse, not AI randomness.

Am I missing something in the link demonstrating that, or is it conjecture?

If you look closely at the response text that accompanies many of these images, you'll find recurring wording like "Here's a diverse image of ... showcasing a variety of ethnicities and genders". The fact that it uses the same wording strongly implies that this is coming out of the prompt used for generation. My bet is that they have a simple classifier for prompts trained to detect whether it requests depiction of a human, and appends "diverse image showcasing a variety of ethnicities and genders" to the prompt the user provided if so. This would totally explain all the images seen so far, as well as the fact that other models don't have this kind of bias.

Re: AI behavior guardrails should be public

#207

Earlier quoted context omitted.

https://pbs.twimg.com/media/GG1eyKjXQAA1FxU?format=jpg&name=... https://cdn.sanity.io/images/cjtc1tnd/production/912b6b5aacc... https://pbs.twimg.com/media/GG1ThfsWUAAp-SO?format=jpg&name=... https://cdn.sanity.io/images/cjtc1tnd/production/e2810c02ff6... https://pbs.twimg.com/media/GG1MnepXwAAkPL6?format=jpg&name=... https://pbs.twimg.com/media/GG0BLVsbMAARZXr?format=jpg&name=...

I don't understand how people could even argue that this is in any way acceptable. Fighting "bias" has become some boogyman and anything "non-white" is now beyond reproach. Shocking.

Seriously, I've basically written off using Gemini for good after this HR style nonsense. It's a shame that Google, who invented much of this tech, is so crippled by their own people's politics.

Re: AI behavior guardrails should be public

#208

How is this any different than doing google image searches of the same prompts. Exmaple: Google image search for "Software Developer" and you get results such that there will be the same amount of women and men event though men make up the large majority of software developers. Had Google not done this with its AI I would be surprised. There's really no problem with the above... If I want male developers in image sea…

> Exmaple: Google image search for "Software Developer" and you get results such that there will be the same amount of women and men event though men make up the large majority of software developers. Now do an image search for "Plumber" and you'll see almost 100% men. Why tweak one profession but not the other?

Because one generates controversy and the other one doesn't.

Re: AI behavior guardrails should be public

#209

Earlier quoted context omitted.

Is there any evidence that this is a consequence of DEI rather than a deeper technical issue?

https://pbs.twimg.com/media/GG1eyKjXQAA1FxU?format=jpg&name=... https://cdn.sanity.io/images/cjtc1tnd/production/912b6b5aacc... https://pbs.twimg.com/media/GG1ThfsWUAAp-SO?format=jpg&name=... https://cdn.sanity.io/images/cjtc1tnd/production/e2810c02ff6... https://pbs.twimg.com/media/GG1MnepXwAAkPL6?format=jpg&name=... https://pbs.twimg.com/media/GG0BLVsbMAARZXr?format=jpg&name=...

"I can't generate white British royalty because they exist, but I can make up black ones" is pretty close to an actually valid reason.

Re: AI behavior guardrails should be public

#210

The gemini guardrails are really frustrating, I've hit them multiple times with very innocuous prompts - ChatGPT is similar but maybe not as bad. I'm hoping they use the feedback to lower the shields a bit but I'm guessing this sadly what we get for the near future.

I use both extensively and I've only hit the GPT guardrails once while I've hit the Gemini guardrails dozens of times. It's insane that a company behind in the marketplace is doing this. I don't know how any company could ever feel confident building on top of Google given their product track record and now their willingness to apply sloppy 'safety' guidelines to their AI.

I had GPT-4 tell me a Soviet joke about Rabinovich (a stereotypical Jewish character of the genre), then refuse to tell a Soviet joke about Stalin because it might "offend people with certain political views".

Bing also has some very heavy-handed censorship. Interestingly, in many cases it "catches itself" after the fact, so you can watch it in real time. Seems to happen half the time if you ask it to "tell me today's news like GLaDOS would".

Post reply on HN