Live data from Hacker News

AI behavior guardrails should be public

twitter.com

261–270 of 350 posts

Re: AI behavior guardrails should be public

#261

Harris and who I think was either Hughes or Stewart a podcast where they talked about how cringey and out of touch the elite are on the topic of race or wokeness in general. This faux pas on google's part couldn't be a better illustration of this. A bunch of wealthy rich tech geeks programming an AI to show racial diversity in what were/are unambiguously not diverse settings. They're just so painfully divorced from r…

Watch your opinion on this get silenced in subtle ways. From gaslighting to thread nerfing to vote locking.... Ask why anyone would engage in those behaviours vs the merit of the arguments and the voice of the people. The strings are revealing themselves so incredibly fast. edit: my first flagged! silence is deafening ^_^. This is achieved by nerfing the thread from public view, then allow the truly caustic to alter…

[flagged]

Re: AI behavior guardrails should be public

#262
post #220

I've never been involved with implementing large-scale moderation or content controls, but it seems pretty standard that underlying automated rules aren't generally public, and I've always assumed this is because there's a kind of necessary "security through obscurity" aspect to them. E.g., publish a word blocklist and people can easily find how to express problematic things using words that aren't on the list. Thing…

This is simply a bad approach and a bad argument. Security through obscurity is a term whose only usage in security circles is derogatory. People figure out how to get around these auto-censors just fine, and not publishing them creates more problems for legitimate users and more plausible deniability for bad policy hidden in them. Doing the same thing but with public policy would already be better, albeit still bad.…

And yet, every single security system on the net relies on an element of obscurity to work. Passwords are secret (obscure), as are private keys for SSL/TLS.

Re: AI behavior guardrails should be public

#263
post #88

Earlier quoted context omitted.

Sure, but this one is from Google adding a tag to make every image of people diverse, not AI randomness.

Am I missing something in the link demonstrating that, or is it conjecture?

Have you bothered to look at all? Read the output of the model when asked about why it has the behaviour it does. Look at the plethora of images it generates that are not just historically inaccurate but absurdly so. It tells you "heres a diverse X" when you ask for X. Yet asking for pictures of Koreans generates only Asian people but prompts for Scots or French people in historical periods generate mostly non-white people. You're being purposefully obtuse, Google has had racism complaints about previous models, talks often about AI safety and avoiding 'bias'. You're trying to argue that it's more likely that the training data had an inherent bias against generating white people in images purely by chance?

Re: AI behavior guardrails should be public

#265
post #142

Earlier quoted context omitted.

I'm convinced this happens because of technical alignment challenges rather than a desire to present 1800s English Kings as non-white. > Use all possible different descents with equal probability. Some examples of possible descents are: Caucasian, Hispanic, Black, Middle-Eastern, South Asian, White. They should all have equal probability. This is OpenAI's system prompt. There is nothing nefarious here, they're asking…

> As these systems get better, they'll figure out that "1800s English" should mean "White with > 99.9% probability". I question the historicity of this figure. Do you have sources?

You're joking surely.

Re: AI behavior guardrails should be public

#266
post #220

Earlier quoted context omitted.

This is simply a bad approach and a bad argument. Security through obscurity is a term whose only usage in security circles is derogatory. People figure out how to get around these auto-censors just fine, and not publishing them creates more problems for legitimate users and more plausible deniability for bad policy hidden in them. Doing the same thing but with public policy would already be better, albeit still bad.…

And yet, every single security system on the net relies on an element of obscurity to work. Passwords are secret (obscure), as are private keys for SSL/TLS.

This misunderstands what is meant by the concept. The mechanisms and the public policies that dictate how they are used are not obscured. To borrow and improve another commenter's analogy about locks and keys, the lock on your door is more secure because every locksmith in the world knows how it works, which doesn't mean they have your key

Re: AI behavior guardrails should be public

#267

Earlier quoted context omitted.

I cannot explain why Google gets a pass, possibly just because they are well entrenched and not an easy target. But AI models are new, they are vulnerable to criticism, and they are absolutely ripe for a group of "antis" to form around.

Well if you have no explanation for that I don’t see why we should try and use your model to understand anything about being risk adverse. They don’t care about being sued, they want to change reality.

That's a pretty unreasonably high standard to hold.

It's an offhand comment in a discussion on the internet not a research paper, expecting me to immediately have an answer to every possible angle here that I haven't immediately considered is a bit much.

Take it or leave it, I don't really care. I was just hoping to have an interesting conversation.

Re: AI behavior guardrails should be public

#268
post #220

I've never been involved with implementing large-scale moderation or content controls, but it seems pretty standard that underlying automated rules aren't generally public, and I've always assumed this is because there's a kind of necessary "security through obscurity" aspect to them. E.g., publish a word blocklist and people can easily find how to express problematic things using words that aren't on the list. Thing…

This is simply a bad approach and a bad argument. Security through obscurity is a term whose only usage in security circles is derogatory. People figure out how to get around these auto-censors just fine, and not publishing them creates more problems for legitimate users and more plausible deniability for bad policy hidden in them. Doing the same thing but with public policy would already be better, albeit still bad.…

Security through obscurity can have a place as part of a larger defense in depth strategy. Alone it's a joke.

Source: in security circles

Re: AI behavior guardrails should be public

#269

I've never been involved with implementing large-scale moderation or content controls, but it seems pretty standard that underlying automated rules aren't generally public, and I've always assumed this is because there's a kind of necessary "security through obscurity" aspect to them. E.g., publish a word blocklist and people can easily find how to express problematic things using words that aren't on the list. Thing…

I think this is a fair approach when things work well enough that a typical user doesn’t need to worry about whether they’ll trigger some kind of special content/moderation logic. If you shadowban spammers and real users almost never get flagged as spammers, the benefits of being tight-lipped outweigh those of the very few users who get improperly flagged or are just curious.

With some of these models the guardrails are so clumsy and forced that I think almost any typical user will notice them. Because they include outright work-refusal it’s a very frustrating UX to have to “discover” the policy for yourself through trial and error.

And because they’re more about brand management than preventing fraud/bad UX for other users, the failure modes are “someone deliberately engineered a way to get objectionable content generated in spite of our policies.” Obviously some kinds of content are objectionable enough for this to be worth it still, but those are mostly in the porn area - if somebody figures out a way to generate an image that’s just not PC, despite all the safety features, shouldn’t that be on them rather than the provider?

Even tuning the model for political correctness is not the end of the world in my opinion, a lot of LLMs do a perfectly reasonable job for my regular use cases. With image generators they are going so far as to obviously (there’s no other way that makes sense) insert diversity sub prompts for some fraction of images which is simply confusing and amateur. Everybody who uses these products just a little bit will notice it. It’s also so cautious that even mild stuff (I tried to do the “now make it even more X” with “American” and it stopped at one iteration) gets caught in the filters. You’re going to find out the policies anyway because they’re so broad an likely to be encountered while using the product innocently - anything a real non-malicious user is likely to get blocked by should be documented.

Re: AI behavior guardrails should be public

#270

Earlier quoted context omitted.

> publish a word blocklist and people can easily find how to express problematic things using words that aren't on the list. I'd love to explore that further. It's not the words that are "problematic" but the ideas, however expressed? Seems like a "problematic" idea, no ?

A word blocklist just serves to apply guardrails. It just slows down common abuse. Very far from perfect, but the alternatives are anything goes or total lockdown. Perfect solutions are pretty damn rare.

You mention "solutions"

Perhaps you mistake me for someone who cares about suggesting solutions for the prolems suffered by giant technopolies, as if they were my problems.

Post reply on HN