Live data from Hacker News

AI behavior guardrails should be public

twitter.com

191–200 of 350 posts

Re: AI behavior guardrails should be public

#191

I've never been involved with implementing large-scale moderation or content controls, but it seems pretty standard that underlying automated rules aren't generally public, and I've always assumed this is because there's a kind of necessary "security through obscurity" aspect to them. E.g., publish a word blocklist and people can easily find how to express problematic things using words that aren't on the list. Thing…

There is no need to implement large scale censorship and moderation in this case. Where is the security concern? That I can generate images of white people in various situations for my five minutes of entertainment? The whole premise of your argument doesn't make sense. I'm talking to a computer, nobody gets hurt. It's like censoring what I write in my notes app vs. what I write on someone's Facebook wall. In one cas…

Our perception of the world has become so abstract that most people can't discern metaphors from the material world anymore

The map is not the territory

Sociologists, anthropologists, philosophers and the like might find a lot of answers by looking into the details of what is included into genAI alignment and trace back the history of why we need each particular alignment

Re: AI behavior guardrails should be public

#192

Earlier quoted context omitted.

This feels more like a personal attack than a response to the argument made.

It's does, but as someone who is staunchly anti-censorship, I understand the frustration. There are sharks out there who want to control speech for their own ends - governments seeking to control populations, corporations wanting docile consumers, hostile nations wishing to stir dissent, individuals trying to cover up their misdeeds, and enabling censorship helps those hostile parties achieve their ends. In this worl…

> better approach would be building up the critical thinking skills of the population

The term for that is "Intellectual self-defence"

> transfers a measure of power to the people and is a multigenerational investment

The challenge to that literacy comes not from power riding on censorship, but from the people themselves who are now conditioned into an apathetic need for "convenience". Critical thinking is hard work, and is never rewarded except in the long run.

Also, anti-censorship absolutism must be tempered with what we call "information hazards". Some of these are genuine, although admittedly very rare, such as easy instructions to create nuclear weapons or synthesise a deadly virus. There are just too many idiots in the world not to want to put a brake on that stuff.

Re: AI behavior guardrails should be public

#193
post #7

I would also love to see more transparency around AI behavior guardrails, but I don't expect that will happen anytime soon. Transparency would make it much easier to circumvent guardrails.

Why is it an issue that you can circumvent the guardrails? I never understood that. The guard rails are there so that innocent people doesn't get bad responses with porn or racism, a user looking for porn or racism getting that doesn't seem to be a big deal.

> The guard rails are there so that innocent people doesn't get bad responses with porn or racism

That seems pretty naive. The "guard rails" are there to ensure that AI is comfortable for PMC people, making it uncomfortable for people who experience differences between races (i.e. working-class people) is a feature not a bug.

Re: AI behavior guardrails should be public

#194
post #12

I strongly suspect Google tried really, really hard here to overcome the criticism is got with previous image recognition models saying that black people looked like gorillas. I am not really sure what I would want out of an image generation system, but I think Google's system probably went too far in trying to incorporate diversity in image generation.

Surely there is a middle ground. "Generate a scene of a group of friends enjoying lunch in the park." -> Totally expect racial and gender diversity in the output. "Generate a scene of 17th century kings of Scotland playing golf." -> The result should not be a bunch of black men and Asian women dressed up as Scottish kings, it should be a bunch of white guys.

It feels like it wouldn't even be that hard to incorporate into LLM instructions (aside from using up tokens), by way of a flowchart like "if no specific historical or societal context is given for the instructions, assume idealized situation X; otherwise, use historical or projected demographic data to do Y, and include a brief explanatory note of demographics if the result would be unexpected for the user". (That last part for situations with genuine but unexpected diversity; for example, historical cowboys tending much more towards non-white people than pop culture would have one believe.)

Of course, now that I've said "it seems obvious" I'm wondering what unexpected technical hurdles there are here that I haven't thought of.

Re: AI behavior guardrails should be public

#195

I've never been involved with implementing large-scale moderation or content controls, but it seems pretty standard that underlying automated rules aren't generally public, and I've always assumed this is because there's a kind of necessary "security through obscurity" aspect to them. E.g., publish a word blocklist and people can easily find how to express problematic things using words that aren't on the list. Thing…

> publish a word blocklist and people can easily find how to express problematic things using words that aren't on the list.

I'd love to explore that further. It's not the words that are "problematic" but the ideas, however expressed?

Seems like a "problematic" idea, no ?

Re: AI behavior guardrails should be public

#196

I've never been involved with implementing large-scale moderation or content controls, but it seems pretty standard that underlying automated rules aren't generally public, and I've always assumed this is because there's a kind of necessary "security through obscurity" aspect to them. E.g., publish a word blocklist and people can easily find how to express problematic things using words that aren't on the list. Thing…

There is no need to implement large scale censorship and moderation in this case. Where is the security concern? That I can generate images of white people in various situations for my five minutes of entertainment? The whole premise of your argument doesn't make sense. I'm talking to a computer, nobody gets hurt. It's like censoring what I write in my notes app vs. what I write on someone's Facebook wall. In one cas…

It's part of the marketing. By saying their models are powerful enough to be gasp Dangerous, they are trying to get people to believe they're insanely capable.

In reality, every model so far has been either a toy, a way of injecting tons of bugs into your code (or circumventing GPL by writing bugs), or a way of justifying laying off the writing staff you already wanted to shit can.

They have a ton of potential and we'll get there soon, but this isn't it.

Re: AI behavior guardrails should be public

#197

I've never been involved with implementing large-scale moderation or content controls, but it seems pretty standard that underlying automated rules aren't generally public, and I've always assumed this is because there's a kind of necessary "security through obscurity" aspect to them. E.g., publish a word blocklist and people can easily find how to express problematic things using words that aren't on the list. Thing…

There is no need to implement large scale censorship and moderation in this case. Where is the security concern? That I can generate images of white people in various situations for my five minutes of entertainment? The whole premise of your argument doesn't make sense. I'm talking to a computer, nobody gets hurt. It's like censoring what I write in my notes app vs. what I write on someone's Facebook wall. In one cas…

What if you are engaged in a wrongthink? How would you suggest this to be controlled instead?

Re: AI behavior guardrails should be public

#198

I've never been involved with implementing large-scale moderation or content controls, but it seems pretty standard that underlying automated rules aren't generally public, and I've always assumed this is because there's a kind of necessary "security through obscurity" aspect to them. E.g., publish a word blocklist and people can easily find how to express problematic things using words that aren't on the list. Thing…

Yes, but the implied problems may not need be approached at all. It's a uniform ideology push, with which people agree differently at different levels. If companies don't want to reveal the full set of measures, they could at least summarize them. I believe even these summaries would be what subj tweet refers to as "ashamed".

We cannot discuss or be aware of the problems-and-approaches, unless they are explicitly stated. Your analogy with content moderation is a little off, because it's not a set of measures that is hidden, but the "forum rules" themselves. One thing is AI refusing with an explanation. That makes it partially useless, but it's their right to do so. Another thing if it silently avoids or directs topics due to these restrictions. Pretty sure authors are unable to clearly separate the two cases, and also maintain the same quality as the raw model.

At the end of the day people will eventually give up and use Chinese AI instead, cause who cares if it refuses to draw CCP people while doing everything else better.

Re: AI behavior guardrails should be public

#199

Earlier quoted context omitted.

> "Generate a scene of 17th century kings of Scotland playing golf." -> The result should not be a bunch of black men and Asian women dressed up as Scottish kings, it should be a bunch of white guys. It works in bing, at least: https://www.bing.com/images/create/a-picture-of-some-17th-ce...

I don't know that this sheds light on anything but I was curious... a picture of some 21st century scottish kings playing golf (all white) https://www.bing.com/images/create/a-picture-of-some-21st-ce... a picture of some 22nd century scottish kings playing golf (all white) https://www.bing.com/images/create/a-picture-of-some-22nd-ce... a picture of some 23rd century scottish kings playing golf (all white) https://www…

> People can explicitly override the sensible defaults as necessary

They cannot, actually. If you look at some of the examples in the Twitter thread and other threads linked from it, Gemini will mostly straight up refuse requests like e.g. "chinese male", and give you a lecture on why you're holding it wrong.

Re: AI behavior guardrails should be public

#200

I've never been involved with implementing large-scale moderation or content controls, but it seems pretty standard that underlying automated rules aren't generally public, and I've always assumed this is because there's a kind of necessary "security through obscurity" aspect to them. E.g., publish a word blocklist and people can easily find how to express problematic things using words that aren't on the list. Thing…

There is no need to implement large scale censorship and moderation in this case. Where is the security concern? That I can generate images of white people in various situations for my five minutes of entertainment? The whole premise of your argument doesn't make sense. I'm talking to a computer, nobody gets hurt. It's like censoring what I write in my notes app vs. what I write on someone's Facebook wall. In one cas…

But what if little timmy asks it how to make a bomb? What if it's racist?

Then what?

Post reply on HN