Live data from Hacker News

AI behavior guardrails should be public

twitter.com

71–80 of 350 posts

Re: AI behavior guardrails should be public

#71
post #7

I would also love to see more transparency around AI behavior guardrails, but I don't expect that will happen anytime soon. Transparency would make it much easier to circumvent guardrails.

Why is it an issue that you can circumvent the guardrails? I never understood that. The guard rails are there so that innocent people doesn't get bad responses with porn or racism, a user looking for porn or racism getting that doesn't seem to be a big deal.

Like a lot of potentially controversial things it comes down to brand risk.

Re: AI behavior guardrails should be public

#72
post #12

I strongly suspect Google tried really, really hard here to overcome the criticism is got with previous image recognition models saying that black people looked like gorillas. I am not really sure what I would want out of an image generation system, but I think Google's system probably went too far in trying to incorporate diversity in image generation.

As well as that, I suspect the major AI companies are fearful of generating images of real people - presumably not wanting to be involved with people generating fake images of "Donald Trump rescuing wildfire victims" or "Donald Trump fighting cops". Their efforts to add diversity would have been a lot more subtle if, when you asked for images of "British Politician" the images were recognisably Rishi Sunak, Liz Truss…

[deleted]

Re: AI behavior guardrails should be public

#73
post #12

I strongly suspect Google tried really, really hard here to overcome the criticism is got with previous image recognition models saying that black people looked like gorillas. I am not really sure what I would want out of an image generation system, but I think Google's system probably went too far in trying to incorporate diversity in image generation.

Remind yourself we're discussing censorship, misinformation, inability to define or source truth and we're concerned on Day 1 about the results of image gen being controlled by a for profit single entity with incentives that focus solely on business and not humanity... Where do we go from here? Things will magically get better on their own? Businesses will align with humanity and morals, not their investors? This is…

> Where do we go from here?

opensource models and training sets. So basically the "secret sauce" minus the hardware. I don't see it happening voluntarily.

Re: AI behavior guardrails should be public

#74
post #73

Earlier quoted context omitted.

Remind yourself we're discussing censorship, misinformation, inability to define or source truth and we're concerned on Day 1 about the results of image gen being controlled by a for profit single entity with incentives that focus solely on business and not humanity... Where do we go from here? Things will magically get better on their own? Businesses will align with humanity and morals, not their investors? This is…

> Where do we go from here? opensource models and training sets. So basically the "secret sauce" minus the hardware. I don't see it happening voluntarily.

Absolutely it won't. We've armed the issue with a supersonic jet engine and we're assuming if we build a slingshot out of pop sticks we'll somehow catch up and knock it off course.

Re: AI behavior guardrails should be public

#75

Earlier quoted context omitted.

>The guard rails are there so that innocent people doesn't get bad responses The guardrails are also there so bad actors can't use the most powerful tools to generate deepfakes, disinformation videos and racist manifestos. That Pandora's box will be open soon when local models run on cell phones and workstations with current datacenter-scale performance. I'm the meantime, they're holding back the tsunami of evil shit…

No legal or financial strategist at OpenAI or Google is going to be worried about buying a couple months or years of fewer deepfakes out in the world as a whole. Their concern is liability and brand. With the opportunity to stake out territory in an extremely promising new market, they don't want their brand associated with anything awkward to defend right now. There may be a few idealist stewards who have the (debat…

Little bit of A, little bit of B.

I am almost certain the federal government is working with these companies to dampen its full power for the public until we get more accustomed to its impact and are more able to search for credible sources of truth.

Re: AI behavior guardrails should be public

#77
post #56

Earlier quoted context omitted.

Is there any evidence that this is a consequence of DEI rather than a deeper technical issue?

You get 4 images per time and are lucky to get one white person when asked for it, no other model has that issue. Other models has no problems generating black people either, so it isn't that other models only generates white people. So either it isn't a technical issue or Google failed to solve a problem everyone else easily solved. The chances of this having nothing to do with DEI is basically 0.

Depending on how broadly you define it, something like 10-30% of the world's population is white. Africa is about 20% of the world population; Asia is 60% of it.

One in four sounds about right?

Re: AI behavior guardrails should be public

#78

The gemini guardrails are really frustrating, I've hit them multiple times with very innocuous prompts - ChatGPT is similar but maybe not as bad. I'm hoping they use the feedback to lower the shields a bit but I'm guessing this sadly what we get for the near future.

I use both extensively and I've only hit the GPT guardrails once while I've hit the Gemini guardrails dozens of times.

It's insane that a company behind in the marketplace is doing this.

I don't know how any company could ever feel confident building on top of Google given their product track record and now their willingness to apply sloppy 'safety' guidelines to their AI.

Re: AI behavior guardrails should be public

#80
post #56

Earlier quoted context omitted.

You get 4 images per time and are lucky to get one white person when asked for it, no other model has that issue. Other models has no problems generating black people either, so it isn't that other models only generates white people. So either it isn't a technical issue or Google failed to solve a problem everyone else easily solved. The chances of this having nothing to do with DEI is basically 0.

Depending on how broadly you define it, something like 10-30% of the world's population is white. Africa is about 20% of the world population; Asia is 60% of it. One in four sounds about right?

It does the same if you ask for pictures of past popes, 1945 German soldiers, etc.
Post reply on HN