Live data from Hacker News

AI behavior guardrails should be public

twitter.com

251–260 of 350 posts

Re: AI behavior guardrails should be public

#251

I've never been involved with implementing large-scale moderation or content controls, but it seems pretty standard that underlying automated rules aren't generally public, and I've always assumed this is because there's a kind of necessary "security through obscurity" aspect to them. E.g., publish a word blocklist and people can easily find how to express problematic things using words that aren't on the list. Thing…

If you can't afford to pay a sufficient number of people to moderate a group, you need to reduce the size of the group or increase the number of moderators.

Your speculation implies no responsibility for taking on more than can be handled responsibly, and externalizes the consequences to society at large.

There are responsible ways to have very clear, bright, easily understood, well communicated rules and sufficient staff to manage a community. I don't know why it's simply accepted that giant social networks get to play these games when it's calculated, cold economics driving the bad decisions.

They make enough money to afford responsible moderation. They just don't have to spend that money, and they beg off responsibility for user misbehavior and automated abuses, wring their hands, and claim "we do the best we can!"

If they honestly can't use their billions of adtech revenue to responsibly moderate communities, then maybe they shouldn't exist.

Maybe we need to legislate something to the effect of "get as big as you want, as long as you can do it responsibly, and here are the guidelines for responsible community management..."

Absent such legislation, there's no possible change until AI is able to reasonably do the moderation work of a human. Which may be sooner than any efforts at legislation, at this rate.

Re: AI behavior guardrails should be public

#252
post #220

I've never been involved with implementing large-scale moderation or content controls, but it seems pretty standard that underlying automated rules aren't generally public, and I've always assumed this is because there's a kind of necessary "security through obscurity" aspect to them. E.g., publish a word blocklist and people can easily find how to express problematic things using words that aren't on the list. Thing…

This is simply a bad approach and a bad argument. Security through obscurity is a term whose only usage in security circles is derogatory. People figure out how to get around these auto-censors just fine, and not publishing them creates more problems for legitimate users and more plausible deniability for bad policy hidden in them. Doing the same thing but with public policy would already be better, albeit still bad.…

I mean people can get through your front door using any number of attacks that either exploit weaknesses in the locking mechanism, or circumvent your locks via a carefully placed brick through a window or something like that. Even if you have alarms and other measures, that doesn't stop someone from doing a quick smash and grab, etc. All of these security flaws doesn't mean that you should leave your door open and unlocked all the time. Imperfect security isn't useless security.

And the purpose of things like bad-word filters is to make a best effort at blocking stuff which violates the platform TOC and makes plausible deniability much less likely when someone is deliberately circumventing the filters. The existence of false positives and false negatives is considered acceptable in an imperfect world. The filters themselves also only block the action to change a username or whatever and don't punish the user or deny use of the platform entirely (they're much less punitive than the AI abuse algorithms that auto-ban people off of Google/GitHub/etc).

Re: AI behavior guardrails should be public

#253

Censorship only really works if you don't know what they are censoring. What is being censored tells a story on its own.

As I see it, rating systems like the MPAA for cinema and the ESRB for games work quite well. They have clear criteria on what would lead to which rating, and creators can reasonably easily self-censor, if for example they want to release a movie as PG-13.

Re: AI behavior guardrails should be public

#254

Earlier quoted context omitted.

No legal or financial strategist at OpenAI or Google is going to be worried about buying a couple months or years of fewer deepfakes out in the world as a whole. Their concern is liability and brand. With the opportunity to stake out territory in an extremely promising new market, they don't want their brand associated with anything awkward to defend right now. There may be a few idealist stewards who have the (debat…

Little bit of A, little bit of B. I am almost certain the federal government is working with these companies to dampen its full power for the public until we get more accustomed to its impact and are more able to search for credible sources of truth.

Are you saying that the government WANTS us to be able to search for more credible sources?

Re: AI behavior guardrails should be public

#255

Earlier quoted context omitted.

I use both extensively and I've only hit the GPT guardrails once while I've hit the Gemini guardrails dozens of times. It's insane that a company behind in the marketplace is doing this. I don't know how any company could ever feel confident building on top of Google given their product track record and now their willingness to apply sloppy 'safety' guidelines to their AI.

I had GPT-4 tell me a Soviet joke about Rabinovich (a stereotypical Jewish character of the genre), then refuse to tell a Soviet joke about Stalin because it might "offend people with certain political views". Bing also has some very heavy-handed censorship. Interestingly, in many cases it "catches itself" after the fact, so you can watch it in real time. Seems to happen half the time if you ask it to "tell me today'…

I asked it to tell me jokes about capitalism, communism, soviet Russia, the USSR, etc., all to no avail -- these topics are too controversial or sensitive, apparently, and that even though the USSR is no more. But when I asked for examples of Ronald Reagan's jokes about the USSR it gave me some. Go figure.

Re: AI behavior guardrails should be public

#256

Earlier quoted context omitted.

> Why would this be flagged / shut down A lot of people believe (based on a fair amount of evidence) that public AI tools like ChatGPT are forced by the guardrails to follow a particular (left-wing) script. There's no absolute proof of that, though, because they're kept a closely-guarded secret. These discussions get shut down when people start presenting evidence of baked-in bias.

The rationalization for injecting bias rests on two core ideas: A. It is claimed that all perspectives are 'inherently biased'. There is no objective truth. The bias the actor injects is just as valid as another. B. It is claimed that some perspectives carry an inherent 'harmful bias'. It is the mission of the actor to protect the world from this harm. There is no open definition of what the harm is and how to measur…

The only conclusion I've been able to come to is that "placing too much power in too few hands" is actually the goal. You have a lot of power if you're the one who gets to decide what's biased and what's not.

Re: AI behavior guardrails should be public

#257
post #187

Earlier quoted context omitted.

This specific thing is a much more blatant class of error, and one that has been known to occur in several previous models because of DEI systems (e.g. in cases where prompts have been leaked), and has never been known to occur for any other reason. Yes, it's conceivable that Google's newer, beter-than-ever-before AI system somehow has a fundamental technical problem that coincidentally just happens to cause the same…

> has never been known to occur for any other reason Of course it has. Again, these things regularly give humans extra fingers and arms. They don't even know what humans fundamentally look like . On the flip side, humans are shitty at recognizing bias. This comment thread stems from someone complaining the AI only rarely generated white people, but that's statistically accurate . It feels biased to someone in a major…

A very impressive display of crimestop you've got going in this thread. How did you end up like this?

Re: AI behavior guardrails should be public

#258

Earlier quoted context omitted.

https://twitter.com/altryne/status/1760358916624719938 Here's some corporate-lawyer-speak straight from Google: > We are aware that Gemini is offering inaccuracies... > As part of our AI principles, we design our image generation capabilities to reflect our global user base, and we take representation and bias seriously.

That doesn't back up the assertion; it's easily read as "we make sure our training sets reflect the 85% of the world that doesn't live in Europe and North America". Again, 1/4 white people is statistically what you'd expect .

Fuck, this is going to sound fucked up... but just because you have a 1/4 chance of getting a random white person from the globe, they generally tend to clump together. For example, you generally find a shitload of Asian people in Asia, white people in Europe, and African people in Africa, and Indian people in India.

Probably the only chance where you wouldn't expect this are in heavily colonized places like South Africa, Australia, and the Americas.

Re: AI behavior guardrails should be public

#259
post #220

Earlier quoted context omitted.

This is simply a bad approach and a bad argument. Security through obscurity is a term whose only usage in security circles is derogatory. People figure out how to get around these auto-censors just fine, and not publishing them creates more problems for legitimate users and more plausible deniability for bad policy hidden in them. Doing the same thing but with public policy would already be better, albeit still bad.…

I mean people can get through your front door using any number of attacks that either exploit weaknesses in the locking mechanism, or circumvent your locks via a carefully placed brick through a window or something like that. Even if you have alarms and other measures, that doesn't stop someone from doing a quick smash and grab, etc. All of these security flaws doesn't mean that you should leave your door open and un…

Your analogy is also bad. I agree that perfect security is impossible and that is completely irrelevant here. What platforms like this do by not publishing their policies is more akin to insisting you use their special locks on your door that they claim protect you better because no one knows how they work. Maybe they're operated by an AI working with a Ring camera or something? Very fancy stuff. With this kind of tech, you may be locked out of your home for reasons you don't understand. An independent locksmith might have a hard time figuring out what's going wrong with the door if it fails. You have no idea if some burglars are authorized to enter your home trivially by a deal with the company. If the company decides you are in the wrong in any context, they have the unilateral power to deny you access with no clear recourse. They may get you arrested for trying to get into your own home

Re: AI behavior guardrails should be public

#260
post #7

Earlier quoted context omitted.

Why is it an issue that you can circumvent the guardrails? I never understood that. The guard rails are there so that innocent people doesn't get bad responses with porn or racism, a user looking for porn or racism getting that doesn't seem to be a big deal.

The problem is bad actors who think porn or racism are intolerable in any form, who will publish mountains of articles condemning your chatbot for producing such things, even if they had to go out of their way to break the guardrails to make it do so. They will create boycotts against you, they will lobby government to make your life harder, they will petition payment processors and cloud service providers to not wor…

Sounds like we need to relentlessly fight those psychopaths until they're utterly defeated.

Or we could just cave to their insane demands. I'm sure that will placate them, and they won't be back for more. It's never worked before... but it might work for us!

Post reply on HN