Live data from Hacker News

AI behavior guardrails should be public

twitter.com

231–240 of 350 posts

Re: AI behavior guardrails should be public

#231
post #100
post #69

Earlier quoted context omitted.

"can you write the rules down so i know them?" --everyone

"No" --Every company that does moderation and spam filtering. "No" --Every company that does not publish their internal business processes. "No" --Every company that does not publish their source code. Honestly I could probably think of tons of other business cases like this, but in the software world outside of open source, the answer is pretty much no.

Then we get back to square one: better no rules at all than secret rules.

This would also be less of a problem if we didn't have a few companies that are economically more powerful than many small countries running everything. At least then I could vote with my feet to go somewhere the rules aren't private.

Re: AI behavior guardrails should be public

#232
post #56

Earlier quoted context omitted.

You get 4 images per time and are lucky to get one white person when asked for it, no other model has that issue. Other models has no problems generating black people either, so it isn't that other models only generates white people. So either it isn't a technical issue or Google failed to solve a problem everyone else easily solved. The chances of this having nothing to do with DEI is basically 0.

Depending on how broadly you define it, something like 10-30% of the world's population is white. Africa is about 20% of the world population; Asia is 60% of it. One in four sounds about right?

If it does, shouldn't there be 60% asians?

Re: AI behavior guardrails should be public

#233

I've never been involved with implementing large-scale moderation or content controls, but it seems pretty standard that underlying automated rules aren't generally public, and I've always assumed this is because there's a kind of necessary "security through obscurity" aspect to them. E.g., publish a word blocklist and people can easily find how to express problematic things using words that aren't on the list. Thing…

Repeat after me: security by obscurity is weak.

Clearly people can work out what some of the rules are, so why not just publish them. If you need to alter them when people figure out how to get around them, well, you already had to anyways.

Re: AI behavior guardrails should be public

#234

Earlier quoted context omitted.

Is a $5 can opener better than a $2000 telescope at opening cans? Yes. Is stable diffusion better at producing finished art, by virtue of not being closed off and DEI'd to oblivion so that it can actually be incorporated into workflows? Emphatically yes. It doesn't matter how fancy your engineering is and how much money you have if you're too stupid to build the right product. As for this being written nonsense, that…

[flagged]

I understood it perfectly.

Re: AI behavior guardrails should be public

#235
post #220

Earlier quoted context omitted.

This is simply a bad approach and a bad argument. Security through obscurity is a term whose only usage in security circles is derogatory. People figure out how to get around these auto-censors just fine, and not publishing them creates more problems for legitimate users and more plausible deniability for bad policy hidden in them. Doing the same thing but with public policy would already be better, albeit still bad.…

Is a content moderation policy the same thing as "security"? Do we get to apply the best practices of the one to the other because they overlap to a smaller or larger degree?

Yes, since content moderation is a form of authorization policy.

Re: AI behavior guardrails should be public

#236
post #220

Earlier quoted context omitted.

This is simply a bad approach and a bad argument. Security through obscurity is a term whose only usage in security circles is derogatory. People figure out how to get around these auto-censors just fine, and not publishing them creates more problems for legitimate users and more plausible deniability for bad policy hidden in them. Doing the same thing but with public policy would already be better, albeit still bad.…

Is a content moderation policy the same thing as "security"? Do we get to apply the best practices of the one to the other because they overlap to a smaller or larger degree?

The use of "security through obscurity" invited the comparison. It is a better comparison when using automated tools instead of human decision-makers. That said, even if we are talking about policy rather than security, policies that are unknown by and hidden from the people they bind is probably the most recognizable and one of the more onerous features of despotism. We have the term "Kafkaesque" because a whole famous writer literally spent his career pointing out how they don't work and harm the people they affect

Re: AI behavior guardrails should be public

#237

Earlier quoted context omitted.

> "Generate a scene of 17th century kings of Scotland playing golf." -> The result should not be a bunch of black men and Asian women dressed up as Scottish kings, it should be a bunch of white guys. It works in bing, at least: https://www.bing.com/images/create/a-picture-of-some-17th-ce...

I don't know that this sheds light on anything but I was curious... a picture of some 21st century scottish kings playing golf (all white) https://www.bing.com/images/create/a-picture-of-some-21st-ce... a picture of some 22nd century scottish kings playing golf (all white) https://www.bing.com/images/create/a-picture-of-some-22nd-ce... a picture of some 23rd century scottish kings playing golf (all white) https://www…

Diversity is cool, but who gets to decide what's diverse?

Re: AI behavior guardrails should be public

#238

I've never been involved with implementing large-scale moderation or content controls, but it seems pretty standard that underlying automated rules aren't generally public, and I've always assumed this is because there's a kind of necessary "security through obscurity" aspect to them. E.g., publish a word blocklist and people can easily find how to express problematic things using words that aren't on the list. Thing…

My dear fellow, some believe the ends justify the means and play games. Read history, have some decency. The danger of being captured by such people far outweighs any other "problematic things" . First and foremost any system must defend against that. You love guardrails so much - put them on the self annointed guard railers. Otherwise, if You Want a Picture of the Future, Imagine a Boot Stamping on a Human Face – fo…

And to add insult to injury the jackbooted thug is an AI bot.

Re: AI behavior guardrails should be public

#239

Earlier quoted context omitted.

https://twitter.com/altryne/status/1760358916624719938 Here's some corporate-lawyer-speak straight from Google: > We are aware that Gemini is offering inaccuracies... > As part of our AI principles, we design our image generation capabilities to reflect our global user base, and we take representation and bias seriously.

That doesn't back up the assertion; it's easily read as "we make sure our training sets reflect the 85% of the world that doesn't live in Europe and North America". Again, 1/4 white people is statistically what you'd expect .

Sure, but I see three 200 responses and a 400 - not 1/4 white people as statistically expected.

Re: AI behavior guardrails should be public

#240

gemini seems to have problems generating white people and honestly this just opens the door for things that are even more racist [1], the harder you try the more you'll fail, just get over the DEI nonsense already 1. https://twitter.com/wagieeacc/status/1760371304425762940

It's not just Gemini, it's Google. And old example is to just search "white people" on Google Images. Almost all the results are black people. https://www.google.com/search?q=white+people&tbm=isch&hl=ro
Post reply on HN