Harris and who I think was either Hughes or Stewart a podcast where they talked about how cringey and out of touch the elite are on the topic of race or wokeness in general. This faux pas on google's part couldn't be a better illustration of this. A bunch of wealthy rich tech geeks programming an AI to show racial diversity in what were/are unambiguously not diverse settings. They're just so painfully divorced from r…
Watch your opinion on this get silenced in subtle ways. From gaslighting to thread nerfing to vote locking.... Ask why anyone would engage in those behaviours vs the merit of the arguments and the voice of the people. The strings are revealing themselves so incredibly fast. edit: my first flagged! silence is deafening ^_^. This is achieved by nerfing the thread from public view, then allow the truly caustic to alter…
AI behavior guardrails should be public
261–270 of 350 posts
Re: AI behavior guardrails should be public
#262I've never been involved with implementing large-scale moderation or content controls, but it seems pretty standard that underlying automated rules aren't generally public, and I've always assumed this is because there's a kind of necessary "security through obscurity" aspect to them. E.g., publish a word blocklist and people can easily find how to express problematic things using words that aren't on the list. Thing…
This is simply a bad approach and a bad argument. Security through obscurity is a term whose only usage in security circles is derogatory. People figure out how to get around these auto-censors just fine, and not publishing them creates more problems for legitimate users and more plausible deniability for bad policy hidden in them. Doing the same thing but with public policy would already be better, albeit still bad.…
Re: AI behavior guardrails should be public
#263Earlier quoted context omitted.
Sure, but this one is from Google adding a tag to make every image of people diverse, not AI randomness.
Am I missing something in the link demonstrating that, or is it conjecture?
Re: AI behavior guardrails should be public
#264Re: AI behavior guardrails should be public
#265Earlier quoted context omitted.
I'm convinced this happens because of technical alignment challenges rather than a desire to present 1800s English Kings as non-white. > Use all possible different descents with equal probability. Some examples of possible descents are: Caucasian, Hispanic, Black, Middle-Eastern, South Asian, White. They should all have equal probability. This is OpenAI's system prompt. There is nothing nefarious here, they're asking…
> As these systems get better, they'll figure out that "1800s English" should mean "White with > 99.9% probability". I question the historicity of this figure. Do you have sources?
Re: AI behavior guardrails should be public
#266Earlier quoted context omitted.
This is simply a bad approach and a bad argument. Security through obscurity is a term whose only usage in security circles is derogatory. People figure out how to get around these auto-censors just fine, and not publishing them creates more problems for legitimate users and more plausible deniability for bad policy hidden in them. Doing the same thing but with public policy would already be better, albeit still bad.…
And yet, every single security system on the net relies on an element of obscurity to work. Passwords are secret (obscure), as are private keys for SSL/TLS.
Re: AI behavior guardrails should be public
#267Earlier quoted context omitted.
I cannot explain why Google gets a pass, possibly just because they are well entrenched and not an easy target. But AI models are new, they are vulnerable to criticism, and they are absolutely ripe for a group of "antis" to form around.
Well if you have no explanation for that I don’t see why we should try and use your model to understand anything about being risk adverse. They don’t care about being sued, they want to change reality.
It's an offhand comment in a discussion on the internet not a research paper, expecting me to immediately have an answer to every possible angle here that I haven't immediately considered is a bit much.
Take it or leave it, I don't really care. I was just hoping to have an interesting conversation.
Re: AI behavior guardrails should be public
#268I've never been involved with implementing large-scale moderation or content controls, but it seems pretty standard that underlying automated rules aren't generally public, and I've always assumed this is because there's a kind of necessary "security through obscurity" aspect to them. E.g., publish a word blocklist and people can easily find how to express problematic things using words that aren't on the list. Thing…
This is simply a bad approach and a bad argument. Security through obscurity is a term whose only usage in security circles is derogatory. People figure out how to get around these auto-censors just fine, and not publishing them creates more problems for legitimate users and more plausible deniability for bad policy hidden in them. Doing the same thing but with public policy would already be better, albeit still bad.…
Source: in security circles
Re: AI behavior guardrails should be public
#269I've never been involved with implementing large-scale moderation or content controls, but it seems pretty standard that underlying automated rules aren't generally public, and I've always assumed this is because there's a kind of necessary "security through obscurity" aspect to them. E.g., publish a word blocklist and people can easily find how to express problematic things using words that aren't on the list. Thing…
With some of these models the guardrails are so clumsy and forced that I think almost any typical user will notice them. Because they include outright work-refusal it’s a very frustrating UX to have to “discover” the policy for yourself through trial and error.
And because they’re more about brand management than preventing fraud/bad UX for other users, the failure modes are “someone deliberately engineered a way to get objectionable content generated in spite of our policies.” Obviously some kinds of content are objectionable enough for this to be worth it still, but those are mostly in the porn area - if somebody figures out a way to generate an image that’s just not PC, despite all the safety features, shouldn’t that be on them rather than the provider?
Even tuning the model for political correctness is not the end of the world in my opinion, a lot of LLMs do a perfectly reasonable job for my regular use cases. With image generators they are going so far as to obviously (there’s no other way that makes sense) insert diversity sub prompts for some fraction of images which is simply confusing and amateur. Everybody who uses these products just a little bit will notice it. It’s also so cautious that even mild stuff (I tried to do the “now make it even more X” with “American” and it stopped at one iteration) gets caught in the filters. You’re going to find out the policies anyway because they’re so broad an likely to be encountered while using the product innocently - anything a real non-malicious user is likely to get blocked by should be documented.
Re: AI behavior guardrails should be public
#270Earlier quoted context omitted.
> publish a word blocklist and people can easily find how to express problematic things using words that aren't on the list. I'd love to explore that further. It's not the words that are "problematic" but the ideas, however expressed? Seems like a "problematic" idea, no ?
A word blocklist just serves to apply guardrails. It just slows down common abuse. Very far from perfect, but the alternatives are anything goes or total lockdown. Perfect solutions are pretty damn rare.
Perhaps you mistake me for someone who cares about suggesting solutions for the prolems suffered by giant technopolies, as if they were my problems.